SpaceX Unveils Grok 4.7 With Enhanced Reasoning and Safety Features
SpaceX has released Grok 4.7, marking its most advanced large language model yet, featuring improved multi-agent processing capabilities and new safety benchmarks.

SpaceX Corp. has unveiled Grok 4.7, representing the company's most sophisticated large language model to date.
The Grok family of models originated from xAI Corp., a startup founded by Elon Musk. Following Musk's merger of xAI Corp. with xAI Inc. in the prior year, the resulting entity was absorbed into SpaceX. Through this acquisition, SpaceX obtained the Grok model line along with several AI data center facilities.
SpaceX assessed Grok 4.7 against CursorBench 4.0, a benchmark created by Cursor, one of SpaceX's recent acquisitions. The model completed benchmark tasks at an average expense of $4.69 per task, surpassing both GPT-5.6 Sol and Fable 5.1. The latter two models were evaluated in hardware-intensive configurations designed to maximize output fidelity rather than minimize operational costs.
Testing with established industry benchmarks revealed mixed results. Grok 4.7 demonstrated substantial improvements over Fable 5.1 when evaluated on the Harvey Legal Agent Benchmark and EEBench, which assess legal reasoning and chip design capabilities respectively. Nevertheless, the model's performance on EEBench trailed behind GPT-6 Astra, OpenAI Group PBC's newest large language model.
Technical Improvements
SpaceX attributes Grok 4.7's capabilities to a newly developed base model. Base models represent the foundational LLM version produced during the initial training phase, before subsequent refinement stages enhance specific algorithmic properties. Many AI developers leverage the same base model across multiple LLM versions.
The company also refined Grok 4.7's reinforcement learning process, a training methodology that strengthens reasoning abilities in base models. Compared to its predecessor, Grok 4.7 underwent more demanding training exercises and spent additional time working through them.
SpaceX engineered Grok 4.7 to integrate with its Grok Bot harness, a framework enabling the model to distribute demanding tasks across multiple AI agents operating simultaneously. This parallel processing architecture accelerates completion times while allowing agents to cross-verify one another's results.
Safety and Pricing
Safety enhancements accompany the release. Grok 4.7 has achieved top scores on LatchBio and HackerBench, evaluation tools that measure LLM resistance to requests for harmful biological research and malicious cybersecurity activities.
Pricing for Grok 4.7 begins at $2 per million input tokens and $6 per million output tokens. SpaceX offers an accelerated variant for time-critical applications that processes requests at double speed, with pricing adjusted accordingly to double the standard rate.
The launch arrives within days of another SpaceX model release. The company introduced Grok Voice Transcribe 2.0, a text-to-speech system, on Friday, delivering twice the accuracy of its predecessor at half the cost. According to SpaceX, the model also outperforms competing text-to-speech offerings in its category.


