Claude Haiku 5.5: what it is and what it costs
2026-10-07 · 4 min read
Claude Haiku 5.5 is Anthropic's new small model, released on October 7, 2026, for fast, high-volume work like summaries, sorting and lookups. It costs $0.10 per million input tokens and $0.50 per million output tokens on prompts under 100,000 tokens, a tenth of what Haiku 4.5 cost.
The launch came as a thread from the Claude account just after 1 p.m. Central. @claudeai called it "the cheapest, fastest, and most capable small model we've ever released," and said it costs around 75% less to run than Haiku 4.5 on average. The opening post passed two million views within hours. It is the last of the Claude 5.5 family, following Opus 5.5 on September 22 and Sonnet 5.5 on September 28.
What Claude Haiku 5.5 costs
The price depends on how long your prompt is. Up to 100,000 tokens, you pay $0.10 per million input tokens and $0.50 per million output, and cache reads drop to a penny per million. Past 100,000 tokens, it jumps to $0.50 in and $2.50 out. Haiku 4.5 charged a flat $1 in and $5 out, so short jobs get 90% cheaper and long ones 50%.
Anthropic leans on the first number. In a follow-up post, @claudeai said prompts under 100,000 tokens make up about 90% of the requests its previous Haiku received.
Simon Willison's launch notes add two caveats. Haiku 5.5 uses a new tokenizer, and his counter tool measured about 1.25 times as many tokens for the same long prompt. You pay less per token but use more of them. Then there's the long-prompt tier. Below 100,000 tokens, Haiku now matches OpenAI's GPT-6 Luna on price. Luna's surcharge doesn't start until 272,000 tokens and only goes to $0.20 and $0.75, so for very large documents he calls Luna "a much better deal."
The same thread also halved cache-read pricing on Sonnet 5.5, to $0.10 per million tokens. Anthropic estimates that cuts Sonnet's cost on most agent tasks by about 20%.
What Haiku 5.5 can do
@claudeai said the model is "a significant step up over Haiku 4.5" in coding, computer use and knowledge work. Anthropic's own numbers back that up by a wide margin. On OSWorld 2.1, a test of operating a real computer, Haiku 5.5 scored 72.4%, where Haiku 4.5 got 15.7% and GPT-6 Luna 48.9%. On GDPval-AA, which grades professional deliverables, it beat Luna by close to 200 Elo points.
It is still the small model, though. Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0 and Haiku 5.5 scored 39.2%, and Anthropic says plainly that Sonnet and Opus are the better choice for complex agentic coding. Haiku is meant to be the helper: the model that compresses a long conversation, summarizes a document or runs quick lookups while a bigger model directs the work.
The effort setting
This is the first Haiku with an adjustable effort level, running from low to max, so each task can lean toward cost or toward quality. Willison found that reasoning can't be turned off entirely and defaults to medium. His standard test, drawing a pelican on a bicycle, took 7 seconds and about a tenth of a cent at low effort. At max effort it took over five minutes and about 3.4 cents. That's cheap either way, but if you multiply it across thousands of daily jobs, the setting you choose will show up on the bill.
Model ID claude-haiku-5-5 is live now on the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure. It returns text only.
What this means for a small business
The repetitive work in a small office suits a model like this: sorting incoming email, pulling details out of a quote request, summarizing call notes, tagging leads in a CRM. Those jobs are short and they repeat all day, and few of them need the smartest model available. At $0.10 per million input tokens the model itself costs next to nothing, so the real expense becomes checking that it got the job right.
If you already run Haiku 4.5 or Sonnet on that kind of work, test Haiku 5.5 at low effort on a week of real jobs before you switch, and count tokens as well as rates, since the new tokenizer changes the math. Keep long-document jobs where they are until you've compared prices at that length.
If you'd like help deciding which tasks belong on a cheap model and which need a bigger one, New Face Design offers a free process audit. We'll map where your hours go and tell you where automation would pay for itself.