Gemini 4 Argon: fewer made-up answers, half-price for now
2026-09-30 · 3 min read
Google finally shipped a new flagship. On Wednesday afternoon @GoogleDeepMind introduced Gemini 4 Argon as "our new frontier model," built for coding, enterprise knowledge work and cybersecurity defense. Seven seconds earlier, Google AI Studio lead @OfficialLoganK posted the price: $2 per million input tokens and $10 per million output tokens "during introductory pricing."
Nobody outside a small group can use it yet. Argon is going first to vetted security teams through Google's Fairwind Program, and Google says it is taking part in the U.S. government's voluntary pre-release testing while access widens. According to Google's announcement, paid API customers and Google AI Ultra subscribers come next.
The independent numbers
Launch posts always look good. The more useful read came about 18 minutes later. @ArtificialAnlys, which runs its own benchmark suite, scored Argon at 53 on its Intelligence Index. That matches GPT-6 Astra and lands one point above GPT-6.1 Sol. For Google it's a big jump: the previous non-Flash model, Gemini 3.1 Pro Preview, scored 30.
The number I'd look at first is about honesty. On the firm's AA-Omniscience test, Argon made things up 15% of the time. GPT-6 Astra did it 51% of the time and GPT-6.1 Sol 54%. Artificial Analysis says Argon is "much more likely to acknowledge when it does not know" rather than guess. The catch is that it got fewer answers right overall, 50% against Astra's 63%. So Argon knows a bit less, and it's more willing to say so.
The arena-style leaderboards agreed. @arena reported that Argon took first place in its Text Arena, where people vote blind between answers, and moved from 29th to 8th in its web development rankings compared with Gemini 3.8 Flash.
The price comes with an asterisk
$2 in and $10 out is the launch discount. Google's own post lists the standard rate as $4 and $20, which is double, and @ArtificialAnlys notes that Google hasn't said when the promotion ends.
The cost math has another wrinkle. Argon writes a lot. Artificial Analysis measured about 62,000 output tokens per task, compared with 27,000 for GPT-6 Astra. At today's discount, a task costs $1.99 on Argon against $3.26 on Astra. At full price, Argon goes to $3.98, which is about 1.2 times Astra. Against GPT-6.1 Sol, which we covered yesterday, Argon already costs 2.7 times as much per task even with the discount.
My read
Google is back in the top tier. Artificial Analysis points out this is its first non-Flash proprietary model in over seven months, so that's been a while coming. Still, the headline benchmark tie matters less to me than the hallucination gap. A model that says "I'm not sure" when it should is much easier to trust with customer-facing work, like answering a pricing question or summarizing a contract. A model that sounds confident and is wrong 51% of the time on hard factual questions needs a human checking its work.
What I'd discount is the price. Introductory rates have a way of ending, and if you build something on $2 and $10 you should run your numbers at $4 and $20 as well. Long answers also add up, so the cost per finished task matters more than the rate card.
Google also says Argon agents found memory optimizations in its own data centers that free up more than 300 TiB once rolled out. That's Google's own claim, and I haven't seen anyone verify it yet.
What this means for a business using AI
Three serious models now sit within a point of each other: Argon, GPT-6 Astra and GPT-6.1 Sol. When the top models are that close, the right pick depends on the work. A model that admits what it doesn't know suits customer questions. For long reports, what matters is the cost per finished task.
Either way, you have to know which tasks you'd hand off first. If you'd like a second set of eyes on that, New Face Design's free process audit looks at where your hours go and which steps an AI model could take over without much risk.