Anthropic's outside referee is Accenture. Anthropic pays the bill
2026-09-19 · 4 min read
Six days after Dario Amodei asked the AI industry to slow down and let outside evaluators in, Anthropic named the first one. It is not who anyone guessed.
The post
On Friday afternoon, @AnthropicAI announced it is "partnering with Accenture on independent evaluation of frontier AI," with both companies expecting to invest "at least $1 billion" in the work over five years. By Saturday morning the post had passed 1.7 million views, and readers had attached a Community Note. The note points out two things the post does not: Anthropic is paying for Accenture's work directly, and the two companies already have a commercial partnership that includes training about 30,000 Accenture staff on Claude.
Both points check out. Anthropic's own announcement says, "Anthropic will fund Accenture's work directly," then adds that "long-term, we think funding should come from pooled or government sources." The commercial tie dates to December, when the companies launched an Accenture Anthropic Business Group and Amodei called the Claude Code rollout there "our largest ever deployment."
What an embedded evaluator is supposed to do
The idea comes straight from Amodei's essay. Instead of testing a finished model from the outside, evaluators sit inside the company with access comparable to an employee's. Per Anthropic, they can watch models during training, follow decisions about what gets built and shipped, talk to staff, and "report incidents and give the public a more informed account of benefits and risks."
Faculty, a UK AI company Accenture bought in January, will run the work. Faculty has tested models for AI labs before, so the expertise claim is not empty. Anthropic says the deal is non-exclusive, that more evaluators will be named in the coming weeks, and that it is talking with METR and other nonprofits about piloting the same setup "using their own funding."
Why the choice landed badly
When Amodei published, the names people expected were research nonprofits like METR, Apollo Research, and Redwood Research. TechCrunch's headline on Friday ended in a question mark for a reason, and it noted Accenture's shares rose about 8% after hours. An evaluator whose parent company makes money reselling and deploying the thing it evaluates is a hard sell as "independent."
Anthropic may have had a reason to go this way. Hours after the essay went up, @DavidSacks, the former White House AI czar who now co-chairs the president's science advisory council, told Anthropic and OpenAI to "go ahead" and slow down, but to "stop pretending METR is independent when it is intertwined with Anthropic's investors and staff." That post has passed nine million views. Picking a large public company that predates the AI boom answers Sacks's objection. It also creates a new one: the referee is now a customer and a channel partner, and the home team is signing the checks.
Anthropic's answer is in the announcement itself: "Independent embedded evaluators do not reduce our accountability, but help to make it more verifiable." I read that as honest. Verifiable is the right word. The X post says independent, and that is the word under strain.
My read
Earlier this week I wrote that step one of Amodei's plan was the only part anyone could check soon. This is step one, and it is checkable. What I want to know is whether Faculty ever publishes a finding Anthropic would rather keep quiet, and whether a release ever slips because of it. TechCrunch points out there are no standards yet for what evaluators can see or say. Anthropic says the same: operational details are still being developed. Until the first uncomfortable report appears, the $1 billion is a press release.
I will give Anthropic this: it disclosed the funding problem in the same document that created it. That is more than most companies do. It also means it knew how this would read and shipped it anyway, presumably because pooled funding does not exist and waiting for it means no evaluator at all.
The question to carry into your own business
Strip away the scale and this is a vendor question every owner faces. Who checked the tool you are about to trust, and who paid them? The consultant who sold you the CRM automation says it works. The agency that built your intake bot says it is accurate. They are probably right, and they are also the last people who should be the only ones looking.
When you evaluate an AI vendor, ask for a reference you can call, not a case study they wrote. Ask who outside their team has reviewed the output. Once something is running, have a person who did not build it read what it sent to customers last week.
New Face Design sells automation, so I will say the quiet part: our free process audit is not independent either. We walk through the work your team repeats and tell you what is worth automating and what is not. You keep the findings whether or not you hire us. Take them to a competitor if you like. That is the only version of independent a small business can buy.