← All posts

Claude Fable 5.1 kept its price. The bill went down anyway

2026-09-02 · 4 min read

On Tuesday afternoon, @claudeai introduced Claude Fable 5.1 and Claude Mythos 5.1 as "the world's most advanced models for coding and knowledge work." The post that matters came four seconds later in the same thread. Cache reads now cost 75 percent less than on Fable 5, and @claudeai says that lowers the real cost "by around 25% for typical workloads, and up to 45% for highly agentic ones."

The list price did not move. Fable 5.1 still charges $10 per million input tokens and $50 per million output tokens, the same as Fable 5. What dropped is the price of a cache read, from $1.00 to $0.25 per million tokens. A cache read is the charge for having the model look again at something it already processed: your instructions, the files you attached, and the conversation so far.

That looks like a footnote until you think about how these models get used now. An agent working through a long job re-reads the same instructions and files on every step, so a job with hundreds of steps is mostly re-reads. Anthropic's launch page has Ramp running the model for 38 hours on one machine-learning problem, and Millennium finding the cause of a crash its engineers had chased for years. On those jobs the meter runs fastest, and it was just cut by three quarters.

The knob most people never touch

Lance Martin (@RLanceMartin) posted migration tips within the hour, and the first is the one to remember. On CursorBench, he says, Fable 5.1 at low effort is "at parity w/ Fable 5 at high effort at a third of the cost." Effort is the setting that decides how long the model thinks before it answers. Claude Code defaults to high, and the Claude app and Cowork default to medium. Anthropic's own docs say the new model matches or beats Fable 5 at low and medium effort and pulls away at the top.

His other tips point the same way. Strip the reassurance out of your prompts: the repeated "double-check your work" lines, the emphasis boosters, the worked examples that went stale two models ago, the rules that contradict each other. Effort can also change partway through a conversation without throwing away the cache, so a hard step gets the full budget while routine steps run cheap.

My read: the sticker price of a frontier model is now the least useful number on the page. The same job on the same model can cost three times as much depending on cache hits and the effort setting. Whoever configures the work sets the bill.

"It fills in the gaps the way I would have"

That line is from Anthropic's Alex Albert (@alexalbert__), who hands the model a few vague, messy sentences and gets back the thing he meant. Anthropic's docs make the same pitch in drier language: the gains concentrate in agentic sessions that run for hours, and it recovers from failed steps when driving a browser or desktop app.

Every's hands-on review shows the other side of gap filling. Asked for 1,000 words, the model returned 1,288. Asked for eight to twelve quotes, it returned 43, and five of the 27 the reviewers could check were not in the source. Anthropic's migration notes list the same tendencies: the model reproduces passages of a source without marking them as quotes, and it rewrites a whole file when a small edit would do.

I would still use it. I would also read "fills in the gaps" literally. When the gap is your intent, the model guesses well. When the gap is a fact, it guesses anyway, and the result reads just as confident.

Mythos 5.1 is the same model with fewer restrictions on cybersecurity and biology work, limited to US organizations in Anthropic's verification programs. For everyone else, Fable 5.1 is the model, and Anthropic says its cybersecurity filter now fires about 60 percent less often per Claude Code session than Fable 5's did.

What this changes for a business

If you use AI through a vendor's product, ask a plain question this week: are they passing the cheaper cache reads on to you, and what effort level are they running your work at? If their cost fell 45 percent and your rate did not, the difference is their new margin.

If you run your own automations, the same two settings decide your bill. Cache the instructions and reference documents every run shares, and run routine steps at low effort. The docs warn that at the lowest effort the model answers from memory more often instead of checking a source, so steps that touch a customer's real data need the higher setting and a check that lives outside the model.

That is the kind of map our free process audit at New Face Design draws: which of your workflows re-read the same material every day and could run at a lower effort, and where a model that fills in gaps should not be trusted to fill them. Tuesday's launch made long AI jobs a lot cheaper to run. It did not make them cheaper to get wrong.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere