Claude Sonnet 5.5: 30% cheaper per task, or 50% pricier?
2026-09-28 · 4 min read
Anthropic released Claude Sonnet 5.5 today. The launch post from @claudeai calls it "a clear upgrade over Sonnet 5," saying it runs more than 30% faster and "costs up to 30% less for most work." @AnthropicAI reposted it a few minutes later.
Within half an hour, the independent benchmarking outfit @ArtificialAnlys posted its own numbers, and they tell a different story about cost. I think both claims are accurate, and working out why is worth a few minutes if you pay an AI bill.
What Anthropic announced
The token prices haven't changed: $2 per million input tokens and $10 per million output, the same as Sonnet 5. Anthropic's announcement page lists 70.6% on Terminal-Bench 4.0, up from 10.3% for Sonnet 5, and 80.1% on the OSWorld 2.1 computer-use test. It says the model comes close to Opus 5.5 on everyday knowledge work. The context window is still 1 million tokens, and the model is available through Anthropic's API, AWS, Google Cloud and Azure.
The "30% cheaper" claim is about the whole task, not tokens. Anthropic's argument is that a faster model making fewer tool calls finishes the same job using fewer tokens overall.
What Artificial Analysis measured
The Artificial Analysis post is long, and mostly positive. Run at max effort, Sonnet 5.5 scored 56 on their Intelligence Index. That puts it second, only 2 points behind Opus 5.5 at max. On several knowledge-work tests (GDPval-AA, AA-Briefcase, AutomationBench-AA) it effectively ties Opus 5.5.
Then there's the catch. At max effort, they say, it produced about 193,000 output tokens per benchmark task, "the highest token use we have measured." That's roughly 60% more than Opus 5.5 or Sonnet 5 at the same setting. Their cost per task came to $7.60, about 50% more than Sonnet 5.
They also found that at lower effort settings, some competing models match its results for less money, and they rate the high setting as its most competitive one. One more caveat: Artificial Analysis tested a pre-release build that had a structured-output bug. Anthropic says that's fixed now, and the firm plans to re-run the affected tests.
Why both numbers hold up
Anthropic is talking about "most work": well-scoped jobs where a quicker model gets it done in fewer steps. Artificial Analysis ran hard benchmark problems at the highest effort setting, where the model keeps thinking as long as it's allowed to. A model can be cheaper on a normal Tuesday and still cost more on a stress test.
For buyers, the price per token doesn't tell you what a job will cost. The effort setting does, along with how many steps the model takes to finish. Sonnet 5.5 comes with five effort levels, from low to max, and choosing one is a budget decision.
The fallback detail
Artificial Analysis also noticed that in about 0.1% of tasks, the model handed the work to Sonnet 5. That's intentional. Sonnet 5.5 is the first Sonnet that comes with Anthropic's cyber safeguards. When a request gets flagged as higher-risk security work, Claude's own apps quietly switch to the older model. According to The New Stack, API users have to opt in to that fallback, so if you just swap the model name in your code, flagged requests may be blocked. That matters if you run an IT or managed-services business. Almost everyone else will never run into it.
My read
This is a strong release. A mid-priced model scoring near the flagship on real office work is good news for anyone automating quotes, reports or inbox triage. But the two posts together are a reminder that launch-day cost claims tell you about the vendor's typical workload, which isn't necessarily yours. I'm glad Artificial Analysis reported cost per task next to the scores, because most launch-day coverage skips it.
What this means for businesses using AI
If you run an automation on Sonnet 5 today, don't switch it over blindly. Run a week of your real jobs on both models, keep the effort setting where your workflow actually needs it, and compare the total cost per finished task. For simple, repeatable jobs, low or medium effort is usually enough, and that's where the savings Anthropic promised should show up.
If you aren't sure which of your processes are worth automating, or what they'd cost to run, New Face Design offers a free process audit. We map where the hours go and tell you honestly whether AI will pay for itself there.