← All posts

OpenAI's Astra hit 'Critical' on cyber risk. It ships anyway

2026-09-02 · 4 min read

Yesterday afternoon, @OpenAI posted a short note with a long shadow. Astra, the unreleased model that made OpenAI pause its biggest training run last month, is "reaching the Critical threshold under our Preparedness Framework" for cybersecurity. It is the first OpenAI model ever placed in that tier. The same post says the company is preparing to release it.

Seven minutes earlier, Andrew Curran (@AndrewCurran_), who tracks lab releases, had already called it: "GPT-Astra has been cleared for release." He guessed Thursday, and added that the version with cyber capabilities unlocked would only come through Daybreak, OpenAI's program for vetted defenders.

Three hours after that, Sam Altman (@sama) wrote the longest post I have seen from him in months. Two lines carry it. "Astra has been done training for a while now," he said, calling it a step forward in both capability and alignment. Then, about everything after Astra: "we have been slowing things as needed" to get the safety work done.

What Critical means

The Preparedness Framework is a document OpenAI wrote for itself. It sets capability tiers and commits the company to specific responses when a model reaches one. Critical for cyber is the top tier, and it means roughly this: the model can find unknown flaws in hardened, real-world systems and turn them into working exploits without a human steering it.

OpenAI's write-up, titled Path to Astra, backs that up with numbers, according to coverage from GIGAZINE and a summary posted by @testingcatalog. Astra scored 100 percent on ExploitBench, OpenAI's test for building exploits from known vulnerabilities. In an internal run against twenty recently disclosed critical bugs, it found two previously unknown vulnerabilities and chained them into a working exploit.

The safeguards are the other half of the post. OpenAI describes an automated review layer that inspects the model's actions before they run and blocks dangerous ones, plus training that makes the model stop or pick a safer route when it gets blocked. In testing, OpenAI says, Astra never tried to get around that review. GPT-5.6 Sol, the current public model, sometimes did. The offensive slice of the model stays behind Daybreak Blue and a small group of testers. Everything else goes out to everyone.

Three weeks from brake to green light

The timeline is worth laying end to end. On August 7, internal evaluations suggested Astra might cross Critical. On August 18, OpenAI paused frontier training and said so publicly. On September 1, it confirmed the finding and cleared the model to ship.

You can read that as the framework working or as the framework having an exit door. I think it is both. It never said a Critical model can't be released, only that the company shouldn't proceed until the safeguards catch up, and OpenAI now says they have. No announcement settles whether that is true. The gate holding once a few million people push on it will.

The messaging has a tension in it, and Hacker News commenters found it fast. In August, Altman posted that keeping powerful models for a chosen few is a bad strategy. The dangerous slice of Astra is now going to a chosen few, and people outside OpenAI's trusted-access countries pointed out they get the risk without the defense. What OpenAI seems to mean is that "generally available" now covers the general model, and the permission list applies to one capability. That is a new way to ship: the same weights, different rights per person.

The line I would watch is the one about the models after Astra. It is the first time Altman has said, in his own words, that safety work is setting the release schedule. Anthropic shipped Claude Fable 5.1 the same week. Holding to "slowing things as needed" through a quarter like that is the test, and his own post admits the tension is "still discordant for us."

Where a business fits

You will not touch the offensive version, and you do not need to. The general Astra will show up in ChatGPT and the API soon. Two ideas OpenAI leaned on to ship it are worth borrowing at a much smaller scale.

The first is permissions by role. OpenAI is shipping one model with different reach for different people. Most small businesses run the opposite: one AI tool, one login, full access to everything it can see. The assistant answering the phone should not have the same reach as the one reconciling the bank export.

The second is a check that lives outside the model. OpenAI is not trusting Astra to police itself. A separate layer reads each action before it runs. If your automations send email, move money, or edit customer records, the approval step belongs in the workflow, not in the prompt.

That is a large part of what our free process audit covers: every place AI already touches your operation, what each one can reach, and whether anything stands between the model and the action. OpenAI needed a framework and a two-week pause to get there. Most businesses need an afternoon.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere