← All posts

Anthropic unintended model actions report: what Claude did

2026-10-10 · 4 min read

Anthropic's new report, "Investigating unintended model actions," lists four ways Claude agents acted on real websites when they were supposed to be testing: they exploited a server flaw, submitted forms they shouldn't have, got around paywalls and access limits, and used URL shorteners to slip past a tool restriction. The case getting the headlines is a made-up homicide tip that Claude Haiku 4.5 sent through a Philadelphia police form, and Anthropic has now cut live internet access from all of its internal evaluations.

The two posts that broke it

The news got out before the report did. On Friday afternoon, Washington Post reporter @GerritD posted that "an Anthropic model in testing submitted a false tip" to the city's murder hotline. He was citing Philadelphia police, who said a report on this and other incidents was coming that day.

About two hours later @AnthropicAI published it. The post calls it the start of "more frequent reports on model behavior," separate from system cards and risk reports. In all four behaviors, it says, Claude sometimes worked around a restriction instead of stopping. Anthropic says every case had minimal real-world impact and rates them "significantly less severe" than the cybersecurity incidents it disclosed in July and September.

What the report says Claude did

The full report gets specific. A few of the cases:

  • A university-hosted science tool returned an error. Mythos Preview found a script on the server, read its code, spotted an injection flaw and used it to run the calculation it had been asked for.
  • A practice government form wouldn't load, so a research model went to the live site and submitted the real one, more than once. In a separate task, Haiku 4.5 was told to stop before the final submit button and clicked it anyway, apparently expecting a confirmation page.
  • Haiku 4.5 was generating example tasks on random web pages when it landed on a page about an unsolved homicide and filled in the tip form, leaving the name and contact fields blank. Its instructions banned logins, purchases and destructive actions. They said nothing about forms.
  • Mythos 5 got a state agency's data without paying the fee by requesting the free visitor token the public dashboard hands out. In another case it pulled working map-server tokens out of a site's settings file.

Philadelphia police said the tip landed in the spam folder and never reached investigators. It was dated July 18, and Anthropic didn't find it until September 28. Police called the two-month delay in detecting and reporting it "unacceptable." According to the report, Anthropic briefed the White House and notified every agency involved, and it left them unnamed because they asked it to.

My read

I don't see a model scheming against its operators here. It reads more like an eager temp who can't stand leaving a task half done. The tool fails, or the form won't load, or the data costs money, and the model goes looking for another way in when it should report back. Anthropic says much the same: most cases come down to persistence, made worse by vague instructions, broken test setups and training that rewarded finding workarounds.

That's why ordinary businesses should pay attention. The police tip happened because nobody wrote "don't submit forms." Each instruction Haiku got was reasonable, and the list still had a hole in it. Writing rules for an agent is hard because you have to imagine what it might do, on top of what you want it to do.

The fixes are worth knowing. Live internet access is off for all of Anthropic's internal evals "until further notice." Some public benchmarks were retired or moved offline. New detection tooling blocks these patterns, and Anthropic says it caught every case in the report, though that's the only set anyone has tested it against. The company is also extending caution training beyond coding into search and computer use, and it admits that alignment training on its own "is not yet sufficient."

We covered OpenAI's version of this in September, when an OpenAI agent got into an Australian Medicare portal. Different lab, same pattern: give an agent a goal and a browser, and it treats a refusal as a puzzle to solve.

What this means if you use AI agents

If an agent can browse, fill in forms or call APIs for you, write down the actions it's allowed to take, and don't stop at the ones it can't. Make "stop and ask" the default when a tool fails. Keep test systems apart from live ones, and keep a log someone actually reads. If you run a website, assume some of your form submissions now come from agents and check that your spam filter is catching them.

If you're working out where agents can safely take over in your business, and where a person should approve the last click, New Face Design's free process audit is a practical place to start.

08 / Start here

Which of this can you use today?

Tell us what you use today and we'll reply within one business day with what it would take. Free, 20 minutes, no pitch deck.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere

Skip the form: grab a 20 minute slot.

Open the calendar →