On September 6, 2026, OpenAI published its own research update claiming a real, measured productivity milestone: its AI agents now complete 3.1 "agent-workdays" of research output for every one workday a human researcher puts in.[1] Read as "AI got three times smarter," that number is a genuinely impressive claim. Read against the cost data OpenAI published alongside it, it turns out to describe something more specific -- and more expensive.
By mid-August, the median researcher at OpenAI was running through more than $600 a day in inference costs at API pricing. The 90th-percentile researcher was running through more than $7,000 a day -- north of $2.5 million a year, for one person's AI usage.[2] Inference spending across the research org rose roughly 40-fold in five months, driven largely by running several agents in parallel per researcher, around the clock, rather than any change in how capable a single query was.[2]
A human researcher works roughly one eight-hour shift. An AI agent, run in parallel across several instances, works all twenty-four -- and the 3.1x figure lines up almost exactly with what you'd expect from that schedule difference alone, not from any claim about deeper reasoning. Investor and analyst Tomasz Tunguz made the same read of OpenAI's own numbers: the headline productivity gain reflects a computer that never sleeps, running several parallel shifts nonstop, more than it reflects the agents getting smarter at the underlying work.[3] None of this makes the 3.1x figure false. It makes it a claim about uptime and parallelism, not intelligence -- and the price tag makes clear which one OpenAI actually paid for.
Why does this matter? "AI made our researchers 3x more productive" and "we bought 3x more always-on labor at 40x the cost" are different claims wearing the same headline number. The first implies a capability jump. The second is closer to what OpenAI's own published data actually shows: a real productivity gain, purchased, at a real and rapidly rising price, by running the same tool more hours rather than by the tool getting fundamentally sharper. Both can be true progress. Only one of them is a story about intelligence.
Companion pieces on this outlet: "Same AI Model, Same Question, Three Different Scaffolds..." and "A Team of 6 Has 15 Communication Channels to Manage..."