← All posts

Musk calls Grok 4.6 objectively No. 1. The scoreboard says third

2026-08-16 · 3 min read

On August 12, @SpaceXAI announced its new flagship model with a plain availability note: Grok 4.6 is "available today in Grok Build, Cursor, Grok Bot, and the API," with double usage in Cursor and Grok Build for the first week. If the handle looks unfamiliar, that's because xAI no longer exists as a separate company. Musk folded it into SpaceX in February, the combined lab has posted as SpaceXAI since July, and Grok kept its name.

A few hours after the launch post, @elonmusk gave his one-line verdict: Grok 4.6 is "objectively #1 when considering intelligence, speed & cost."

That same day, @ArtificialAnlys, the independent benchmarking firm whose Intelligence Index has become the default scoreboard for frontier models, posted its own number. Grok 4.6 scores 61, "joining the frontier in line with GPT-5.6 Sol." Here is the top of the index with that score in place:

  • Claude Opus 5 at maximum reasoning: 63
  • Claude Fable 5: 62
  • Grok 4.6 and GPT-5.6 Sol at maximum reasoning: 61
  • Kimi K3: just behind

Third place, two points off the lead. So was Musk wrong?

The case for Musk's math

Not entirely, and that's the interesting part. In a follow-up post, Artificial Analysis pointed out that Grok 4.6's headline pricing didn't move from Grok 4.5: still $2 per million input tokens and $6 per million output. The five-point intelligence gain came at a cost per task comparable to Kimi K3, an open-weights model, and far below what Claude Opus 5, GPT-5.6 Sol, or Claude Fable 5 cost to run. If your ranking is capability per dollar, Grok 4.6 has a genuine claim. If your ranking is raw capability, it's third.

Musk chose the ranking that flatters his model, the way every vendor does. What matters for the rest of us is that both rankings are now defensible, because the raw-capability race has gotten absurdly tight.

Two points now separate four labs

Look at that list again. Anthropic, OpenAI, and SpaceXAI sit within two points of each other, with an open-weights model from Moonshot close behind. And the pace is as striking as the spread: Grok 4.5 shipped just over a month before 4.6 and scored five points lower. At that cadence, this month's standings tell you almost nothing about next quarter's.

There's a tell in the launch post itself. Double usage for the first week is promotional pricing, the kind of offer a streaming service runs, or a gym in January. You run promotions when the products are close enough that customers might actually switch, and frontier models have reached that point. The fight has moved from "our model can do things yours can't" to price, speed, and how long the model can grind on a task without supervision, which is exactly what SpaceXAI says 4.6 is tuned for.

If you buy AI instead of benchmarking it

For a small business wiring AI into quoting, follow-ups, or reporting, the leaderboard question has quietly stopped mattering. Whichever frontier model you pick this month will be roughly matched, and possibly passed, within a quarter. The questions that actually move your costs are different ones. What does a finished task cost, not a million tokens? Does the model handle your specific workload well? And can you swap it out without rebuilding everything when the standings flip again?

That last one is where businesses get stuck. An automation hard-wired to one vendor in 2025 is a renegotiation problem in 2026. We build ours at New Face Design so the model is a swappable part, which turns a price war like this one into leverage instead of lock-in. If you're weighing where AI would actually pay off in your operation, our free process audit is a low-stakes place to start.

The scoreboard will probably reshuffle again before the promo pricing expires. The systems you build around it don't have to.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere