← All posts

OpenAI's agents log 3.1 workdays for every human one

2026-09-08 · 4 min read

On Sunday, OpenAI researcher Kevin Liu (@kliu128) posted that the company was releasing data on how its own models are speeding up its own research. His stated reason was not the numbers. It was that this kind of progress happens inside a handful of labs and nobody outside gets to see it, so "being transparent is more urgent than ever," and he asked other AI companies to publish the same.

The headline figure is a ratio. Measured in eight-hour workdays, OpenAI's research organization now runs about 3.1 agent-workdays of effort for every workday a human researcher puts in. That was mid-August. Before June, total agent runtime across the research org was still below total human labor, so the crossover happened this summer.

By mid-August the median OpenAI researcher was burning more than $600 a day on coding-agent inference at API prices, and the top tenth were past $7,000 a day. The company also says it hit a goal Sam Altman set last fall: an automated "research intern," meaning a system that can take a well-defined task worth a few days of a skilled person's time and carry it out under human direction. Next target on the board is a fully automated AI researcher by March 2028.

Noam Brown (@polynoamial) called it "one of the most interesting blog posts we've released" and pointed at the second half of it, which covers how OpenAI has paced model development around monitoring, alignment, and security.

The number that did not trend

Buried in the same post: over the previous six months, more than half of the four to eight hour agent tasks that succeeded still needed at least one human intervention along the way. That is the success column. The failures are a separate accounting.

What that describes is an agent taking a brief, running for most of a workday, and getting steered back on course at least once by someone who already knows roughly what the answer should look like. OpenAI is fairly straight about this in the post. The researcher still defines the question, checks the output, and owns whether it was right.

What 3.1 actually measures

Concurrency, mostly. A researcher who kicks off four agents before lunch is banking machine-hours while eating a sandwich, and those agents spawn subagents of their own. The ratio tracks how much machine work is in flight, not how much more gets finished.

That distinction got lost fast. On Monday, @MRRydon took the 3.1 figure, drew a straight line from a 1:1 ratio on June 1, and landed on 10:1 by late April 2027, concluding that "things are going to go vertical very fast." It is a clean extrapolation through four months of data on a metric that counts runtime, and runtime is cheap to add. You add it by opening more tabs.

My honest read: the underlying shift is real and the disclosure is genuinely useful, because almost nobody publishes this. But the lab is grading its own homework with metrics it picked, and those daily dollar figures are API prices for compute OpenAI already owns. So 3.1 is a fair description of how the work now feels inside the lab. It is not a productivity multiplier anyone outside has checked.

What this looks like at your scale

Strip out the frontier-lab context and the working pattern is the same one that pays off in a ten-person company. Nobody wins by handing an agent one enormous vague job. You win by running several bounded jobs at once and checking in on them.

That is the practical takeaway from the intervention number. If more than half of the long agent tasks at OpenAI still need a human course-correction, the automation you build should have the checkpoints designed in rather than bolted on after something goes sideways. A quote draft that waits for approval before it sends. A weekly report that flags the two things it was unsure about instead of quietly guessing.

Notice also that someone at OpenAI is metering this per researcher, per day. Cost per task has become a line item at the company least likely to care about it. For everyone else it should be the first question you ask about an automation, well before anyone books a demo.

If you are trying to work out which parts of your operation would actually hold up to this, that is what our free process audit is for. We look at how work moves through your business and tell you where an agent would help and where it would just create a new thing to check.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere