Grok 4.7 promises stronger AI coding performance without a higher base price, but heavy token consumption could reduce those savings.
SpaceXAI released Grok 4.7 on Sept. 21 as its latest model for coding and professional knowledge work. According to the company, it uses a larger base model than Grok 4.6 and underwent a longer reinforcement-learning run focused on difficult tasks that can take hours to complete.
The training is designed to improve how Grok verifies its own work, handles long contexts and manages multi-step assignments. SpaceXAI also trained it to natively understand the Grok Bot harness, which the company says improves conversational performance and general knowledge work.
Coding gains come with caveats
In company-reported testing, Grok 4.7 outperformed Grok 4.6 across several coding benchmarks. On CursorBench 4.0, Grok 4.7 scored 46.3%, up from 40.4%. Its Terminal-Bench 4.0 score nearly doubled to 38%, compared with 20.3% for its predecessor.
It also reached 71% on DeepSWE v1.1 at high effort and 64% on EEBench, compared with 65.2% and 53%, respectively, for Grok 4.6.
But the gains do not amount to a clean sweep. SpaceXAI‘s own comparison puts Fable 5.1 Max ahead of Grok 4.7 on CursorBench and Terminal-Bench. Grok also trails competing models on some professional knowledge-work and clinical benchmarks in SpaceXAI’s comparison. Independent testing raises another concern. Artificial Analysis reported that Grok 4.7 at its xHigh reasoning setting used about 81,000 output tokens per Intelligence Index task, substantially more than Grok 4.6 High and GPT-6 Astra Max in its testing.
The price is the real hook
Grok 4.7 starts at $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6. The faster version costs $4 per million input tokens and $12 per million output tokens.
That pricing makes the model attractive for developers running large workloads, but the token consumption complicates the calculation. A cheaper token does not automatically produce a cheaper completed task if the model needs substantially more tokens to finish it. For companies, the useful measurement will therefore be cost per successful job, not simply the API price.
New safeguards for cybersecurity
SpaceXAI says Grok 4.7 introduces a new safeguard stack designed to improve jailbreak resistance while limiting unnecessary refusals.
The company reports a 62.4% score on LatchBio’s biosafety benchmark and says its safeguards allowed 3.3% of the risky dual-use prompts tested with HackerBench v0.3. Selected cybersecurity partners are also receiving invite-only access to Grok 4.7 for red-team testing.
Those figures come from SpaceXAI’s own testing, so broader independent evaluations will be important.
What Grok 4.7 means for users
For developers, Grok 4.7 is most relevant when work involves repositories, terminals, repeated tool calls and long coding sessions rather than simple prompts. Professionals could also benefit from its focus on documents, presentations and extended knowledge-work tasks.
The model is available through Cursor, Grok Build, the Grok API, third-party coding tools, model routers and cloud platforms.
The main limitation is predictability. Benchmark results vary by methodology, effort setting and toolchain, while high token consumption could raise real-world costs.
Where the release could matter most
Grok 4.7’s unchanged pricing could make it attractive for agentic AI workflows in which models operate for extended periods rather than answer individual prompts. However, organizations should test the model using their own repositories, tools and approval processes before assuming its lower API rates will reduce overall spending.
The more useful measurement will be cost per successful task, including token consumption, retries, latency and human review. Independent evaluations and production results will ultimately determine whether Grok 4.7’s benchmark gains translate into dependable savings.