← All posts

GPT-6 Sol costs half as much and says 'I don't know' more often

2026-09-22 · 4 min read

About 90 minutes after Anthropic shipped Claude Opus 5.5 on Tuesday, OpenAI answered with two models of its own. The @OpenAI account introduced GPT-6 Sol and GPT-6 Luna as faster, cheaper models that carry over much of GPT-6 Astra's strength, with "50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing."

So both big labs spent the same Tuesday cutting prices. The question for anyone paying the bill is what the lower price buys, and the most useful answer so far came from a firm that benchmarks models for a living.

The price and the pitch

Sol is the heavy model, aimed at coding and agent work. It now costs $2 per million input tokens and $10 per million output, down from $4 and $20. Luna is the budget tier at $0.10 and $0.50, down from $0.20 and $1.20. OpenAI pitches Luna for "summarizing documents, extracting information, or answering quick questions," which covers a lot of the paperwork in a small office.

OpenAI's own headline claim is about mistakes. On an internal test built from real conversations where users flagged errors, it says GPT-6 Sol makes about half as many mistakes as the model it replaces. Both models are rolling out in ChatGPT Work and Codex for paid plans and in the API, and Free and Go users can try Luna in the desktop app.

The independent check: cheaper, not smarter

Eight minutes after the launch post, @ArtificialAnlys posted its results with a blunt summary: "Intelligence Index and Coding Agent Index scores remain level with GPT-5.6." This is a price cut. Capability stayed where it was.

The savings hold up, though. Running Artificial Analysis's full test suite cost $1.06 per task on GPT-6 Sol, against $1.99 on GPT-5.6 Sol. Luna went from $0.18 to $0.07. Both new models write slightly more per task than their predecessors, so all of the gain comes from the lower rate.

Two other findings in that post matter more to a business than the headline scores.

The first is about made-up answers. On Artificial Analysis's knowledge benchmark, Sol's hallucination rate fell from 92% to 60%, mostly because it declined to answer more often: it attempted 83% of the questions, down from 99%. Wrong answers dropped by about a quarter, while overall accuracy slipped from 59% to 54%. Sol didn't learn more facts. It got better at saying "I don't know."

The second is that office deliverables got worse. On GDPval-AA, a test of real professional tasks across 44 occupations, Sol lost about 100 Elo points and Luna about 75. The testers inspected hundreds of outputs by hand and traced the drop to weaker presentation and deliverables that left out required pieces.

A user's read

@danshipper, whose team at Every had tested Sol before launch, called it "my new daily driver in Codex," adding that it's "not quite Astra, but close enough" for much of his everyday work. He also ran it head to head against Opus 5.5. In his tests Sol writes clean prose that leads with the point, while Opus 5.5 has the higher ceiling on long autonomous coding builds.

He had one real complaint. Codex's new security classifier kept stopping work he had already approved so it could ask for approval again. Shipper blames the classifier rather than Sol, but that kind of friction is often what decides whether a team keeps using an agent at all.

One detail stood out to me. Anthropic says Opus 5.5 puts the most important information first, yet Shipper found its drafts still bury the point. Launch claims and user tests disagree all the time, which is a good reason to run your own.

My read

For anyone paying per token, this is a good release. For anyone hoping for a smarter model, it's a boring one. I'd pay the most attention to the shift toward declining rather than guessing.

When a model is reading invoices, answering customer questions, or pulling fields out of forms, "I don't know" is the answer you want when it isn't sure, because a confident wrong answer costs far more than a blank. The catch is that blanks need somewhere to go. If your workflow assumes the model always returns something, a more cautious model will quietly leave gaps.

What it means for a business using AI

  • Luna's price makes bulk clerical work close to free at the token level. The real cost is checking the output, so plan for that check from day one.
  • Give every automation a fallback. When the model declines, the item should land in front of a person instead of disappearing.
  • Look hard at anything a client will see. The GDPval drop suggests polished reports and proposals from these models may need more human review, not less.

Most of the work in a reliable automation is picking the right model for each step and deciding what happens when it doesn't know. If you want help mapping that out for your own processes, our free process audit is a good place to start.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere