OpenAI says it has reached the “automated research intern” milestone it set for itself, and it has put two numbers behind the claim. Its researchers are now running 3.1 agent-workdays for every human workday, and the heaviest external users are spending upward of $7,000 a day on tokens.
Agent-workdays is a new unit, and it is a smart one to publish. It sidesteps benchmark arguments entirely and reports something closer to economic substance: how much autonomous work is running in parallel with the humans directing it. A ratio above three means each researcher is supervising more machine effort than they could personally perform.
What it does not tell you
Parallelism is not productivity. Three agent-days of work that go nowhere still count as three agent-days. Without a measure of what fraction of that output survives review and lands in something shipped, the ratio describes throughput of attempts rather than yield. Internal metrics announced by the company that invented the metric deserve that caveat.
The $7,000-a-day figure is the more interesting disclosure, because it is a market fact rather than an internal one. Someone is paying that voluntarily and repeatedly, which puts a floor under the argument that agent workloads are a demo rather than a business line.
Why the framing matters
“Research intern” is a deliberately modest label for a system that is being described as doing multiples of a person’s daily work. An intern needs direction, produces work that gets checked, and occasionally gets things badly wrong. That is roughly the right expectation, and it is a more honest one than most of the marketing in this category.
The next threshold is the one that would actually matter: agent output that a senior researcher accepts without rewriting it. Nobody has claimed that yet.