A free AI model now writes exploits. Patch faster
2026-09-30 · 4 min read
On Tuesday, Anthropic's Frontier Red Team published a report on GLM-5.3, the newest open-weight model from the Chinese lab Zhipu AI, which goes by Z.ai outside China. The finding is uncomfortable: a model anyone can download now builds working cyber exploits nearly as well as Claude Mythos Preview, the model Anthropic decided was too risky to release to the public.
The reactions on X split quickly. Security researcher @lukOlejnik put the headline numbers side by side: "Mythos for vetted users, GLM-5.3 free for all." He also pointed out that the report "reads like a product sheet," since Anthropic's conclusion is that defenders need models as strong as the attackers' models, and Anthropic sells those models. Researcher @willcb made the same point as a joke, saying Anthropic's message amounted to "glm-5.3 is evil and chinese."
I think both jabs are fair. They still don't change the numbers.
What the report actually found
Anthropic's write-up tested GLM-5.3 on ExploitBench, a set of tasks where the model has to turn a known software flaw into a working attack. GLM-5.3 succeeded on 50 of 410 attempts, about 12%. Mythos Preview managed 56, about 14%. On a harder internal benchmark covering full control-flow hijacks, GLM-5.3 hit 4% against Mythos at 6%. The team's point there is that earlier models, including Claude Opus 4.6 and GLM-5.2, succeeded on none of those tasks.
Twelve percent sounds low until you remember that an attacker can just run the model again. Runs are cheap, and one success is enough.
The cost figures are the part a business owner should care about. Olejnik's summary says turning a public Chrome bug into a working exploit cost $20.40 in model fees. Anthropic's report describes a known flaw becoming a working attack with about 20 minutes of human effort, with the model doing the rest.
Then there are the guardrails. GLM-5.3 ships with some refusal training, but Anthropic says simple tricks got past it most of the time: 64% with deceptive prompts, 92% with prefilled reasoning. Stripping the safety behavior out of the weights entirely, a technique called abliteration, worked every time and cost roughly $1,200 to $4,400 in compute. Once the weights are public, nobody can take that back.
My read
The commercial angle is real. Anthropic has every reason to argue that closed, supervised models are safer, and this report helps that case. I'd keep the incentive in mind without throwing out the results. The benchmarks are specific, and the comparison is against Anthropic's own best model, which is not a flattering yardstick to pick if you only wanted to scare people.
The open-versus-closed fight will keep going on X. For a business, the practical change is simpler. The gap between a vulnerability going public and someone holding a working attack for it is getting shorter, and per this report it can now be closed with 20 minutes of human attention and about twenty dollars.
What this means for a small business
Most small businesses don't get breached by novel attacks. They get breached through known holes that nobody patched, like an old WordPress plugin, a router that hasn't had a firmware update since it was installed, or a remote-access tool with a default password. Cheap, automated exploit writing makes those known holes much easier to use.
So the defensive work hasn't changed much. It just has less slack now:
- Know what you run. You can't patch software you forgot you have, so keep a plain list of your site plugins, your devices, and your logins.
- Turn on automatic updates wherever you can, and set a weekly check for the systems where you can't.
- Put multi-factor authentication on email, banking, and anything with admin access.
- Keep backups you have actually tried restoring, and store one copy offline.
You don't need a frontier model for any of that. You need a routine, and a lot of it can run on its own: update checks, plugin audits, an alert when something falls behind. If you aren't sure where your gaps are, New Face Design's free process audit is a reasonable place to start. We look at how your business actually runs day to day, and that includes the boring maintenance work.