← All posts

Grok 4.7 stopped giving up early. Its token use more than doubled

2026-09-21 · 4 min read

Ten days ago, @elonmusk explained why Grok 4.7 had missed its launch date. The model "still gives up on hard tasks (that it can do!) too early," he wrote, and it wasn't checking its own work carefully enough. His guess at the cause was that reinforcement learning had penalized long responses too much. Training had pushed the model to keep answers short, and on hard problems short meant unfinished.

The fixed version shipped today. SpaceXAI's launch page calls Grok 4.7 its most powerful model for coding and knowledge work, built on a larger base model with longer reinforcement learning on harder tasks. Pricing stays at $2 per million input tokens and $6 per million output, same as Grok 4.6, and it's live in Cursor, Grok Build, and the API.

The same day, @ArtificialAnlys posted independent numbers showing that the fix worked, and what it cost.

What the scoreboard says

Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index, two points above Grok 4.6. The firm says that brings SpaceXAI into the top four AI labs. It isn't a new leader, though: The Decoder reports Claude Fable 5.1 and GPT-6 at 53 each on the same index.

The bigger gains are in the long, grinding work Musk said the model was quitting on. On the firm's Coding Agent Index, Grok 4.7 running in Grok Build scores 56, up nine points, which ranks fourth among models in their own harnesses behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. On AA-Briefcase, a private benchmark for long-horizon knowledge work, it lands next to Opus 5 and Fable 5.1. On everything else, the firm says it broadly matches Grok 4.6, with small slips on a couple of tests.

That's close to Musk's own pre-launch grade. In a Sept. 14 reply, he put 4.7 "roughly on par with Opus 5.0" and said Grok 4.9 is "probably Astra/Fable class." Founders rarely forecast their own model that modestly, and the independent data mostly backs him up.

The line that matters

The scores got the headlines, but the line in the Artificial Analysis thread I'd pay attention to is this one: "Grok 4.7's gains come with higher token usage."

At its highest reasoning setting, Grok 4.7 used about 81,000 output tokens per index task. Grok 4.6 used about 36,000, and GPT-6 Astra at its maximum setting used about 27,000. That's 125% more than its predecessor and nearly three times OpenAI's flagship.

Read that next to Musk's diagnosis. The old model stopped too soon, the new one keeps working, and you pay for every token it spends doing that. The price per token didn't change, so the cost of a finished task went up.

Two caveats, to be fair about it. Artificial Analysis tested 4.7 at "xhigh" effort and 4.6 at "high," so some of the jump comes from the setting rather than the model. Reasoning effort also runs from low to xhigh, so you can turn the thinking down. The public numbers don't yet show how much of the gain you keep when you do.

SpaceXAI's pitch is "half the price of comparable models," which is a per-token claim. Suppose a rival really does charge twice as much per token but finishes the same job in a third of the tokens. Then the rival is the cheaper way to get the job done.

What it means for a business buying AI

The frontier labs keep offering the same trade, more thinking for better results. For some work that's the right call. A model reconciling a messy month of job costs should grind through it. A model sorting inbound leads into three buckets shouldn't be thinking for long at all.

If you run AI inside quoting, intake, or reporting, a few habits follow from this:

  • Track the cost of a finished task, not the price per million tokens. Your invoice counts tokens, and this upgrade uses a lot more of them.
  • Set reasoning effort per job instead of leaving everything on maximum.
  • Re-test when you switch models, since a "same price" upgrade can double what a task costs.

Before we recommend a model for a workflow at New Face Design, we measure what each step costs to complete and check which steps need a heavy thinker in the first place. If you want that picture for your own operation, our free process audit is a simple place to start.

Musk was right that the old model quit too early. The version that doesn't quit costs more to run, and whether that's worth it depends on the job you give it.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere