← All posts

OpenAI's GPT-6 Astra hides more of its thinking

2026-09-04 · 4 min read

OpenAI started rolling out GPT-6 Astra on Wednesday, and the launch language was not subtle. Greg Brockman reportedly closed the briefing with a line about the AGI era arriving, and @gdb followed up on X saying he is excited to see what it does for entrepreneurship and for "how small teams tackle big problems."

Set the AGI argument aside. The number worth your attention this week came out of the safety evaluations, not the keynote.

@scaling01 pulled it from the system card: the UK AI Safety Institute clocked Astra's time horizon at "30.9 minutes compared to 3.6 minutes for GPT 5.6 Sol." That is the no chain-of-thought horizon, a measure of how much human-equivalent work the model gets through in a single pass without writing down its reasoning along the way. Roughly an order of magnitude, in one release.

Ryan Greenblatt, chief scientist at Redwood Research, read the same launch and objected. He wrote that Astra reportedly uses an architecture where "more of the reasoning occurs in activations instead of natural language," and called it "the single worst development for AI security/safety to date".

What changed under the hood

Reasoning models became trustworthy partly by narrating. The model wrote out its steps, and that written trail turned out to be useful twice over: it made the answers better, and it let researchers read what the thing was doing on the way there. Chain-of-thought monitoring is built on that side effect.

Astra reportedly leans on a technique called recurrent depth. Instead of pushing a query forward through the network once, it loops the same layers over the same query several times. You get more thinking per parameter, which is a real efficiency win. The loop runs in numbers rather than sentences, so the extra thinking leaves no written trace.

OpenAI pushed back on the alarm. Per TechCrunch, chief scientist Jakub Pachocki said the company has worked to preserve chain-of-thought monitoring since its first reasoning models, and that Astra's computational depth stays within roughly twice GPT-4's, with the looping capped on purpose.

Greenblatt's reply to that is the sharpest thing said all week. He allowed that the transparency is welcome, then noted the statement is also consistent with Astra having a dial that is currently set low and could be turned up.

My honest read

Both things are true. Recurrent depth is good engineering, and the safety worry is not hypothetical. A capped loop today is a capped loop by choice, and choices made under competitive pressure tend to drift in whichever direction wins benchmarks. Anthropic and Google DeepMind are reported to be weighing the same technique, which is how a race to the bottom on transparency usually starts.

For most people using ChatGPT this is invisible. For anyone building on top of these models, the useful takeaway is narrower and more practical: the reasoning trace you have been reading was never a guarantee, and it just got less complete. Treating it as an audit trail was always a little generous. Now it is measurably less true.

What this means if you are putting AI into your business

Every serious automation I build ends up sitting on the same question: when this goes wrong, how will you know, and what will you be able to show?

The model does not answer that question. What you build around the model does. Logging the input and output of every AI step covers most of it, along with deciding up front which calls the system is allowed to make by itself and where a person still has to sign off, which should include anything that moves money or goes out under your name. None of that depends on the reasoning being legible, which is why it survives the reasoning becoming less legible.

An owner who wired AI in with the model as the only record will feel this. An owner who kept receipts at every step will not notice it happened. Models change every few weeks now, so build it so you can swap the model out and still have the trail.

If you want a read on where that applies in your own operation, a free process audit at New Face Design walks your real workflows and marks the steps where an AI decision needs a witness. Sometimes the answer is that a process should stay manual, and that is a fine outcome too.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere