5 دقيقة قراءة
Grok 4.7 Is Cheaper Than the Frontier and Sharper in a Few Places: Where It Fits for AI Agents

By submitting, you consent to our use of your data. Privacy Policy.
Category
وكلاء الذكاء الاصطناعي
Share the article
The most useful way to read a Grok 4.7 launch is to ignore the leaderboard position, because it is not the best model, and xAI has not really pretended otherwise. What Grok 4.7 is, is cheap for what it does, unusually strong in a few specific domains, and built for the kind of long-running work agents actually do. For an enterprise deciding where to point which model, that is a more interesting profile than another try at the top spot.
xAI released Grok 4.7 on 21 September as its most capable model for coding and knowledge work, priced the same as Grok 4.6 and shipping the same day into Cursor, GitHub Copilot, the Grok API, and the major model routers. The headline is not a benchmark. It is the price next to the benchmark.
What Grok 4.7 is
Grok 4.7 is a reasoning model that xAI trained with a longer reinforcement-learning run on a harder mix of tasks, deliberately weighted toward problems that take many hours to finish. That focus on long-horizon work shows up in what the model is built to do: manage extended context, and check its own work as it goes rather than at the end.
It costs $2 per million input tokens and $6 per million output, the same as the previous version, with a fast variant that doubles the output speed at twice the price. That pricing is the whole strategy, and it is worth its own section.
The benchmarks, read honestly
Grok 4.7 is not topping the charts, and pretending otherwise would not survive contact with the numbers. On general agentic coding it trails the frontier: 46.3% on CursorBench 4.0 against Claude Fable 5.1 Max's 51.8%, and 38.0% on Terminal-Bench 4.0 against Fable's 57.9%. On the hardest long-running agentic work, Astra and Fable still look stronger overall.
Where it gets interesting is the specialised domains, where the gains are large and sometimes decisive.
Benchmark | Grok 4.7 | Best comparison |
|---|---|---|
CursorBench 4.0 (agentic coding) | 46.3% | Fable 5.1 Max 51.8% |
Terminal-Bench 4.0 | 38.0% | Fable 5.1 Max 57.9% |
DeepSWE v1.1 (high effort) | 71.0% | competitive |
EEBench (electrical engineering) | 64.0% | Grok 4.6 was 53.0% |
Harvey Legal Agent Benchmark | 19.6% | GPT-5.6 Sol Max 2.5% |
That legal number is not a typo. On the Harvey legal-agent benchmark Grok 4.7 scored more than seven times what the comparison frontier model managed, and it jumped eleven points on electrical engineering in a single version. The pattern is a model that is average-to-good on general tasks and genuinely differentiated on specific ones.
The real story is the price
xAI held the price flat and let the rest of the market move. At $2 and $6 per million tokens, Grok 4.7's output pricing undercuts other frontier models by somewhere between three and eight times. For a chatbot that barely matters. For an agent it matters enormously, because an agent runs long chains of reasoning, and output tokens are where a long-running agent's bill actually accumulates.
That is the calculation an enterprise should run. A model that is a few points behind the frontier on general work, but a fraction of the cost per output token, wins the moment the work is high-volume or long-running, which describes most of a back office.
Where Grok 4.7 fits in an agent stack
The move, as with every launch this year, is not to standardise on it. It is to route the right work to it. Grok 4.7 earns a specific and useful set of steps.
Point it at the cost-sensitive, long-running, and tool-heavy work where its price and its multi-hour training pay off, and at the specialised domains, legal and engineering among them, where it is now genuinely differentiated. Keep the hardest general reasoning and the most demanding long-horizon agentic tasks on Astra or Fable, which still lead there. This is the same logic as the open versus closed decision: the question is not which model is best, it is which model belongs on which step.
What early users are saying about Grok 4.7
In the first day after launch, the reaction split along the same line the benchmarks draw. The consistent praise is for cost and speed, which matches the pricing story. The consistent question is whether it is actually a step up in capability, with more than one person asking the obvious thing: does 4.7 solve problems 4.6 could not, or is it mostly faster and cheaper.
There is also healthy scrutiny of the "checks its own work" claim, with developers asking what the self-verification actually verifies, and some early complaints about usage limits in the tools it shipped into. None of that is a verdict. It is the sensible reading for any launch: the specialised and cost numbers are real, and whether the general capability moved is something to confirm on your own tasks over the next few weeks, not from a launch chart.
Common questions about Grok 4.7
Is Grok 4.7 better than GPT-6 Astra or Claude Fable 5.1?
Not overall. On general agentic coding and the hardest long-running tasks, Astra and Fable still lead, with Fable ahead of Grok 4.7 on CursorBench and Terminal-Bench. Grok 4.7's advantages are price and specific domains like legal and electrical engineering, where it is competitive or ahead.
How much does Grok 4.7 cost?
$2 per million input tokens and $6 per million output, the same as Grok 4.6, with a fast variant at double the speed and double the price. Its output pricing undercuts other frontier models by roughly three to eight times, which is its main advantage for long-running agents.
Is Grok 4.7 good for AI agents?
For the right steps, yes. Its low output price and training on multi-hour tasks suit cost-sensitive, long-running, and tool-heavy agent work, and it is strong in specialised domains. Route the hardest general reasoning to a higher-scoring model and measure Grok 4.7's completion rate on your own jobs before committing.
Where can I use Grok 4.7?
It launched into Cursor, GitHub Copilot, Grok Build, the Grok API, and the major model routers on day one. Advanced red-team cybersecurity capabilities are gated to invite-only partners.





