← All posts

Claude now leads 26% of building the next Claude. Claude graded it

2026-09-18 · 4 min read

On Thursday, @AnthropicAI posted that AI systems are getting more powerful and "they're increasingly being used to build the next version of themselves." The company wants "to illuminate that progress for the public." The post introduced three measurements: how much AI research is done by AI, how well the agents doing it are overseen, and where the compute goes. Then it published a snapshot of all three from inside Anthropic.

The number that traveled: as of August, Claude "leads" 26% of Anthropic's AI research and development work. In February it was under 1%.

What "leads" means

The full report rates tasks on an automation scale built by Epoch AI. It runs from AL0, no AI involved, to AL5, AI working with no human in the loop. Two levels matter here. At AL3, "collaborates," Claude does large chunks of the work under close human direction and a person approves the decisions. At AL4, "leads," Claude carries most of a task from a high-level prompt to the finish while a human supervises without watching every step.

So 26% of the work is at AL4, more than 90% is at AL3 or higher, and nothing measured is at AL5. Anthropic says it plainly: Claude "is not operating fully autonomously for any measured subset of AI R&D work."

The tasks themselves are ordinary engineering. The report's task tree has 542 nodes, and the examples it gives include diagnosing defects in the evaluation platform, network policy for training sandboxes, and writing postmortems after serving incidents.

The other two numbers

About 30,000 agents were working at any one time on Anthropic's main internal platform in August. An online monitor screens every action before it runs, usually within seconds. It exists for harms that can't be undone, and the report's example is an agent copying model weights out of Anthropic's systems. Of more than a billion decisions that month, it blocked about 1 in 47,000. A slower offline monitor flags one or two transcripts per thousand actions for human review.

The compute figure covers one week in July. That week, 6% of AI R&D compute went to safety work. Inside the AI-led slice, it was 12%.

Researcher Elie Bakouch, quoting the announcement, read the two labs' disclosures side by side. @eliebakouch noted that OpenAI published usage per R&D task while Anthropic published automation levels, both drawn from the same Epoch AI framework, and that it "would be great to combine both." He also flagged that "only 6% of research compute went to ai safety." OpenAI's numbers came out earlier this month. We covered them in the 3.1 workdays post.

My read

The caveats are the most honest part of the report, and Anthropic wrote them itself.

The ratings were produced mostly by Claude. When staff rated the same tasks, they matched the model exactly 59% of the time and landed within one level 97% of the time. The task list was frozen, so the 26% describes automation of work that already existed, not new categories of work Claude invented. The compute number covers one week on one platform. And the line between safety research and capabilities research is, in Anthropic's own words, subjective.

That leaves the 26% as a self-reported baseline rather than a wrong number. The post says any frontier lab could publish the same measures "and third parties could verify them." Nobody outside has yet. What Anthropic has put on the table is a public scale and a public method, with an open invitation to audit the results. That follows through on the pacing essay Dario Amodei published last week, which promised embedded evaluators. The first outside number is the one I'd wait for.

What this means for your business

The scale is the part an owner can use today. Every automated task in your company sits somewhere between AL0 and AL5, whether you have named it or not. A drafted reply that a person sends is AL3. A bot that answers the phone and books the appointment on its own is AL4. Most businesses I talk to want AL4 for internal reports and AL3 for anything a customer sees, and they have never written that down.

The monitor is the other lesson. Anthropic runs 30,000 agents with a checker on every action, and it still blocks about one in 47,000. Whatever you automate, decide what the checker is for you. For anything that pays or sends, that means an approval queue. For the rest, a daily log that someone reads.

New Face Design's free process audit puts a level next to each of your workflows and names who checks it. Start here.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere