← All posts

Google's new Gemini isn't smarter. That's the whole point

2026-07-22 · 4 min read

On Monday, @GoogleDeepMind announced "three new models to make AI agents faster, smarter, and cheaper at scale": Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a security-tuned Gemini 3.5 Flash Cyber. A follow-up post said 3.6 Flash "builds directly on feedback from 3.5 Flash," with a demo comparing quality and token usage side by side.

Then the independent scoreboard weighed in. @ArtificialAnlys posted its verdict within hours: 3.6 Flash "maintains the same Intelligence as Gemini 3.5 Flash." Same score, 50, on their Intelligence Index. On a raw IQ leaderboard, Google's newest model moved exactly zero places.

That sounds like a miss. I think it is the most honest release of the summer.

What actually shipped

The headline numbers are not about intelligence at all:

  • Gemini 3.6 Flash uses about 17 percent fewer output tokens than 3.5 Flash to do the same work, and output pricing drops from 9 dollars to 7.50 per million tokens, with input at 1.50.
  • Artificial Analysis clocked the average agentic task at 1.3 minutes, down from 2.7. Same intelligence, half the wait.
  • The knowledge cutoff jumps from January 2025 to March 2026, and coding got a real bump: Google cites 49 percent on DeepSWE versus 37 for the old model.
  • Flash-Lite lands at 30 cents per million input tokens for high-volume grunt work, and Flash Cyber, which finds and patches software vulnerabilities, stays in a locked pilot for governments and trusted partners only.

Meanwhile the elephant in the room did not ship. Gemini 3.5 Pro, the flagship, has now missed multiple targets, and Google's line is that it arrives "as soon as it's ready." In the same breath, the company said it has already started its most ambitious pretraining run yet, for Gemini 4. Translation: please look past the delayed model at the bigger one behind it.

The race quietly changed lanes

Here is my read. For two years, every model launch was a claim about being smarter. This launch is a claim about being cheaper and faster at the same smartness, and Google barely pretended otherwise. The pitch in the announcement is agents at scale, not genius in a box.

That shift matters more than another two points on a benchmark. When AI runs as an agent, it does not answer one question and stop. It takes a task, reasons through steps, calls tools, retries, and burns tokens the whole way. At that point three things decide your bill and your user experience: tokens per task, seconds per task, and price per token. Gemini 3.6 Flash improved all three and left the IQ dial alone.

The community reaction split along exactly that line. People staring at the Intelligence Index saw a flat score and shrugged. People who run agents in production saw a model that finishes the same job in half the time for less money and did the math.

What this means if you run a business

You do not need a frontier genius to answer your phones, chase unpaid invoices, follow up on quotes, or fill canceled appointments. Those jobs were within reach of last year's models. What kept many small businesses on the sidelines was the operational side: cost per run, speed, reliability at volume. That is precisely the front where the labs are now competing, and prices are falling while task times shrink.

It also means the smart move is rarely "wait for the next big model." The flagship slipped again. The workhorse tier got better on the metrics that hit your ledger. If an automation made marginal sense at last quarter's prices, it probably clears the bar today, and it will clear it more comfortably in six months.

The practical question is not which model tops the leaderboard. It is which of your repeatable processes leak the most time and money, because those are the ones this efficiency race keeps making cheaper to fix. If you want a straight answer on where AI would pay for itself in your operation, New Face Design offers a free process audit for businesses here in the Fox Valley and beyond. We will map your workflows and show you the two or three where a boring, fast, cheap model quietly earns its keep.

Google spent Monday bragging that its new model thinks the same but wastes less. For once, the marketing and the reality point the same direction.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week actually goes and identify the first process worth automating. You keep the map either way. No pitch deck, no pressure.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere