Gemini 3.8 Flash: same price per token, 40% more per task
2026-09-02 · 4 min read
On Wednesday morning, @GoogleDeepMind announced two models at once: Gemini 3.8 Flash, "our most intelligent model yet with significant gains from 3.7 Flash," and a locked-down Gemini 3.8 Flash Cyber for finding and patching software vulnerabilities. Gemini 3.7 Flash came out on August 13. That is a three-week gap between versions.
The independent scoreboards posted within minutes. @ArtificialAnlys called it Google's "fourth Flash model in under four months" and scored it 59 on its Intelligence Index, three points above 3.7 Flash and level with GPT-5.6 Sol and Grok 4.6 when those run below their maximum reasoning settings. For scale, Claude Fable 5.1 sits at 66 on the same index and Claude Opus 5 at 63. Flash is not the smartest model available. It is now within reach of models that cost many times more.
Where it landed
@arena ranked it in three arenas the same hour. In Agent Arena, 3.8 Flash debuted at number 14, up from 32 for 3.7 Flash. In Text Arena it took seventh place, two points ahead of Claude Opus 5. In the web development coding arena it landed at 18.
Google's own launch page picks three benchmarks where 3.8 Flash edges Opus 5: a finance agent test, Harvey's legal agent test, and HLE-Verified, where the gap is half a point. The Wall Street Journal reported ahead of launch that Google engineers testing the model on the company's internal Jetski coding platform preferred it to Opus.
My read: the improvement is real, and the "beats Opus" framing is thin. The wins are on tests Google chose, by margins inside the noise, and the internal preference came from Google staff on a Google tool. The fair headline is duller. A model priced like a budget option now sits near the top ten on independent boards, and it got there three weeks after the previous version.
The line owners should read twice
Gemini 3.8 Flash keeps the introductory price of 3.7 Flash: 75 cents per million input tokens and $3.75 per million output tokens, through December 31. Google lists the standard rate at double that.
Same price per token does not mean the same bill. Artificial Analysis measured 3.8 Flash at 58 cents per task on its index, about 40 percent more than 3.7 Flash. The reason is that the new model writes more. Average output per task rose 30 percent to around 48,000 tokens, and it takes more turns on agent-style evaluations. Google describes this as the point. In the launch post, Google's Tulsee Doshi and Raluca Ada Popa write that 3.8 Flash "works harder" on complex tasks, "executing extra reasoning steps, and calling tools iteratively."
Put this next to the rest of the week. On Tuesday, Anthropic held Claude Fable 5.1 at the same list price and cut the cost of cache reads by 75 percent, so long jobs got cheaper. On Wednesday, Google held its price and shipped a model that spends more tokens per job, so long jobs got pricier and better. In July, Google's own 3.6 Flash went the other direction and cut tokens per task by 17 percent. Neither lab is wrong. Each release makes the sticker price per million tokens a worse guide to your monthly bill.
Four models, four months, one budget
The release pace matters more than any one score. If your business wired something to Gemini 3.5 Flash in May, you are three versions behind today, and each version changed how the model behaves as well as how it scores. A model that reasons longer and calls tools more often will finish jobs the old one failed. It will also take longer and cost more per run, and sometimes it will do more than you asked.
- Budget on the standard rate, not the intro rate. The discount ends December 31, and the standard price is twice the number in the announcement.
- Track cost per completed job, not cost per token. That is the number that moved 40 percent this week while the price sheet stayed flat.
- Re-test before you switch versions. Run the new model on last month's real tickets or documents and compare output length and error rate first.
- Pin the version you tested. Google's release cadence means "latest" changes every few weeks.
The Cyber variant deserves one note. Google is gating it behind its Fairwind Program for governments, critical infrastructure operators, and software maintainers. That mirrors what Anthropic did with Mythos 5.1 on Tuesday and what OpenAI is doing with Astra. All three labs now hand their best vulnerability-finding model to a vetted list instead of the public.
If you run a small business, the Opus comparison is not your problem. Your problem is whether a model this cheap can now handle the document-heavy work that clogs your week, and whether you would notice if its cost per job drifted up. That is what our free process audit at New Face Design maps: which of your workflows is worth automating at today's model prices, and what to measure so the next three-week release does not surprise your bill.
Google spent Wednesday saying its new model works harder. The scoreboard and the meter both agree.