Grok 4.7 arrived on September 21, 2026. For teams buying AI to write code or analyze business information, its release poses a practical question: does a better model make finished work cheaper? The answer depends on more than its advertised token rates. Official release notes
That mixed picture is the subject of Matthew Berman’s September 22 video, I don’t know how to feel about Grok 4.7…. He questions the choice of comparisons and argues that cost per completed task matters more than token prices. He also says he has not tested the model thoroughly. His video is an early assessment of the evidence, not a comprehensive hands-on evaluation. Berman’s introduction, 0:00
The launch deserves attention because a useful model does not have to win every benchmark to change what businesses can afford to automate. It does, however, have to finish enough work correctly to justify its total cost. Independent testing already exposes the tension: Grok 4.7’s better results can come with much more generated output.
For developers, small businesses and teams building agents, the decision is therefore practical. Which tasks can this model complete at an acceptable quality? How much effort does it need? And how much work remains for a person after the model says it is done?
We reviewed the complete captions of Berman’s roughly 17-minute video and checked the launch, product documentation and Artificial Analysis’s evaluation. The analysis below includes no BIG CHANGE benchmark or claim of first-hand Grok testing.
What launched, and where it is available
The standard model is available on the xAI API as grok-4.7, as well as in Cursor and Grok Build. The API accepts text and images and produces text. Its documented context window is 500,000 tokens, with four reasoning settings: low, medium, high and xhigh. High is the default. Official release notes
SpaceXAI attributes the improvement to a larger base model and longer reinforcement learning on harder tasks. This is the developer’s explanation. Launch technical overview
There is also Grok 4.7 Fast. The documentation describes it as the same model on faster infrastructure, billed at twice the standard token rates. It is offered through Cursor and Grok Build, is excluded from Grok Build’s free tier and is not available through the public xAI API. Grok 4.7 product documentation
Those distinctions matter when comparing screenshots, demonstrations and bills. Model version, reasoning effort and serving tier are separate choices. A demonstration running Fast at xhigh is not an estimate of the experience or cost of a standard-speed, medium-effort API call.
The launch page’s comparison with other models should also be kept separate from the upgrade comparison. SpaceXAI says standard Grok 4.7 has the same price and speed as Grok 4.6. It does not say that upgrading from 4.6 automatically halves a customer’s bill.
The price card has a long-context threshold
The published standard API rates are:
- Up to 200,000 prompt tokens: $2 input, $0.50 cached input and $6 output per million tokens.
- Above 200,000 prompt tokens: $4 input, $1 cached input and $12 output per million tokens.
These are token rates, not all-inclusive prices for an agent completing a job. The model page flags higher pricing for requests exceeding 200,000 tokens, and the release notes specify the higher rates. Pricing and availability were checked on September 22. Model pricing and release notes
A 500,000-token context window therefore does not mean every request within it uses the lowest rate. Nor does a large context window establish that the model will consistently find every relevant detail in a large collection of files.
Cursor introduces another product-level distinction: its documentation lists a 256,000-token standard window and a 500,000-token maximum. The capacity exposed in an editor can differ from the API specification or depend on the chosen mode. Cursor’s Grok 4.7 documentation
For a team processing long reports, the immediate question is whether sending all the material improves the result enough to justify its cost. Retrieving the relevant sections may be cheaper and easier to inspect. Sometimes the full document is necessary; sometimes it is simply the easiest way to construct a prompt. Those situations should not be budgeted as if they were identical.
Better performance can consume more tokens
Artificial Analysis’s September 21 evaluation gives Grok 4.7 at xhigh a score of 46 on its Intelligence Index, two points above Grok 4.6. It reports approximately 81,000 output tokens per task, compared with 36,000 for Grok 4.6 at high effort. The same report gives 38,000 for Grok 4.6 at xhigh, so even that same-effort comparison shows more than twice as many output tokens. Artificial Analysis’s launch evaluation
This is a concrete reason to separate token price from task cost. Using 81,000 output tokens at the base $6-per-million rate produces an illustrative output charge of $0.486. At 38,000 tokens, it would be $0.228. Those calculations deliberately exclude input, cache behavior, tool charges, higher-context pricing and failed attempts. They are arithmetic using the reported token counts, not measured all-in job prices.
More output is not automatically wasteful. A model might spend extra effort checking an answer and avoid a failure that would otherwise require a person to intervene. It might also spend longer pursuing a wrong approach. The token count alone cannot distinguish productive persistence from expensive repetition.
Berman returns to this problem near the end of his video, where he notes the higher token usage reported by Artificial Analysis. Video, 15:43
The useful result is an accepted deliverable. A code change that passes appropriate tests and review, a spreadsheet whose formulas reconcile, or a report whose claims trace to evidence has a different value from an answer that merely arrives quickly.
The benchmarks show a capable but uneven model
The launch table gives Grok 4.7 xhigh 46.3% on CursorBench 4.0, below Fable 5.1 max at 51.8%. Fable also leads its Terminal-Bench comparison. Vendor benchmark table
The accompanying chart plots benchmark score against output tokens, decreasing toward the right. Its lines compare reasoning settings. Original interactive chart: select Tokens. Open the full-size chart.

Moving upward means a higher score; moving rightward means fewer generated tokens in this evaluation. It does not mean a lower all-in price: token rates, input usage and the surrounding agent still matter. This is also a different test from the Artificial Analysis token figures above. Combining their counts would erase the task and methodology differences that make either result interpretable.
That is enough to reject a blanket claim that Grok 4.7 leads every kind of coding work. It is also enough to justify a closer look for particular workloads. The size of an improvement on one evaluation does not determine how a model will handle an unfamiliar repository, an incomplete bug report or an unusual tool environment.
Artificial Analysis provides a useful independent distinction. Its report gives Grok 4.7 with Grok Build a Coding Agent Index score of 56, up from 47 for Grok 4.6 with that agent. It explicitly separates these native-agent results from its standardized Intelligence Index setup. Independent evaluation and methodology notes
The surrounding agent matters. It supplies tools, manages context and decides how the model’s requests are executed. Changing that environment can change results even when the model name stays the same. Combining a terminal score from one setup with a software-engineering score from another can create a persuasive table that describes no system anyone actually tested.
Berman flags the absence of GPT-6 Astra from one launch table and uses an AI-generated reconstruction to add it. That is useful as a prompt to investigate the comparison; the reconstruction is not an independent benchmark source. We do not adopt its added figures without checking the underlying evaluation. Video, 6:03
There is a commercial relationship to keep in view as well. Cursor’s August 14 company announcement states that it was acquired by SpaceX. A benchmark published within that product family remains useful evidence, but it is not a disinterested assessment of a competing supplier. Cursor’s acquisition announcement
Office work may be the more interesting opportunity
The independent evaluation reports an AA-Briefcase score of 1,657 Elo for Grok 4.7 at xhigh, 111 above Grok 4.6 at high. It attributes much of the improvement to analytical quality, while presentation quality is slightly lower than the earlier model’s result. Artificial Analysis’s knowledge-work results
That split is useful for anyone commissioning reports, spreadsheets or presentations. An answer can contain stronger analysis and still need editing or redesign. Conversely, a polished deck can conceal shallow reasoning. Evaluating only the visible finish risks rewarding the wrong improvement.
An operations team could test the model on a recurring performance report. A useful trial would check whether it uses the correct period, reconciles figures with the source data and distinguishes an observed change from a proposed explanation. The person reviewing it should record how much correction was needed, not just whether the report looked professional on first inspection.
A smaller business might care less about winning the most difficult benchmark and more about making an otherwise unaffordable piece of analysis routine. That is a plausible route to broader adoption. It becomes real when the output is accurate enough, arrives on time and leaves less work for the owner than doing the job manually.
Professional benchmark labels should not be mistaken for professional qualifications. A legal-work or clinical-reasoning result describes performance on that evaluation. It does not establish that a model can independently handle a client’s legal matter or make a patient-care decision. This article makes no recommendation to use Grok for those decisions.
Safeguards are another claim to evaluate in context
SpaceXAI says Grok 4.7 has a new safeguard stack and stronger jailbreak resistance. That is a launch claim, not a guarantee that every application using the model will be safe. The developer’s safety account
For a business agent, at least two questions remain. Does the model respond appropriately to the requests it receives? And does the application limit what the model can actually do?
A model can correctly understand an instruction to change a customer record while lacking permission to make that change. It can also encounter malicious instructions embedded in a document it was supposed to summarize. Access controls and approval rules must remain part of the surrounding system; improved refusal behavior does not replace them.
The practical implication is to test the whole workflow. Include legitimate requests that should succeed, disallowed actions that should be blocked and ambiguous cases that should pause for clarification. A product that refuses harmless work too often can be unusable; one that executes unauthorized actions can be much worse. Neither issue is captured by a single headline score.
What the video leaves unresolved
Berman’s coverage raises several questions that should remain questions. His explanation linking pricing to available compute is a market interpretation, not disclosed unit-cost accounting. The social-media predictions he discusses are expectations, rather than evidence of delivered performance. His closing demo comparison comes with his own acknowledgment that he does not know the settings used. Pricing discussion, 8:44, expectations, 11:51 and demo discussion, 15:20
The video also discusses open-weight models. That broader competitive trend should not be confused with Grok 4.7’s licensing: Artificial Analysis identifies this release as proprietary and lists its weights as unavailable. Grok 4.7 release profile
These distinctions keep the evaluation focused. A product can be attractively priced without the public knowing why its supplier chose that price. A striking failed demonstration can identify a test worth repeating without establishing that the model is generally poor. An available model can be useful without matching a founder’s earlier forecast.
There is also a sponsorship boundary. Berman’s video includes a Zapier advertisement. It is not independent evidence of Grok’s performance, and its promotional claims are not recommendations from BIG CHANGE.
A better buying metric: cost per accepted task

A team considering Grok 4.7 should compare it with its current process on a small, representative set of jobs. Define acceptance before seeing the outputs. For coding, that might include relevant tests, a readable patch and no unrelated changes. For analysis, it might include correct calculations and evidence for consequential claims.
Then record the complete bill: input and output usage, tools, retries and the time spent reviewing or repairing the result. Count unsuccessful attempts too. Divide that cost by the number of deliverables that meet the acceptance criteria.
An illustrative example shows why the distinction matters. Suppose one configuration costs $20 across a batch and produces 80 accepted results. Its measured cost is $0.25 per accepted result before human labor. Another costs $12 but produces only 30 accepted results, giving $0.40 per accepted result. The cheaper batch is the more expensive source of usable work. These are hypothetical numbers, not Grok measurements.
Reasoning effort should be part of that trial. A high-effort setting might earn its cost on a difficult task and add little on a routine one. Escalating only the cases that need it could be more economical than running every request at xhigh, provided the routing decision itself is evaluated.
Keep review standards constant across models. Lowering the bar for a cheaper model creates an apparent saving by accepting worse work. Raising the bar only for the incumbent creates the opposite distortion. The point is to buy a dependable outcome, with a clear account of what people still have to do.
The change worth watching
Grok 4.7 gives teams another option to test. Choosing it requires a view of the entire job: what the model produces, what it consumes and what someone still has to repair.
The broader consequence could be more selective use of AI across everyday work. A team may choose one configuration for routine analysis, another for difficult coding and a person for cases where judgment or accountability cannot be delegated. More competition gives that team another option to test; it does not dictate which option should win.
For BIG CHANGE, the next milestone is measurable work completed with less total effort. If Grok 4.7 reduces the cost of accepted deliverables, it can make useful automation available to more organizations. If extra reasoning merely moves expense from the token price into the volume consumed, the headline price will tell buyers very little. The evidence that settles that question will come from completed jobs, including the ones that failed first.



