- Work Shift: OpenAI says coding agents logged 3.1 eight-hour runtime units for every human workday across its research organization by mid-August 2026.
- Human Control: People still set priorities and judge results; over half of successful four-to-eight-hour tasks received an intervention.
- Evidence Limit: Code and experiment activity increased during 2026, but the figures do not demonstrate gains in productivity, quality, model capability or scientific novelty.
OpenAI says its coding agents were logging 3.1 eight-hour agent-workdays for every human workday across its research organization by mid-August 2026. The agents execute delegated software tasks, offering a rare view of how AI is redistributing work inside a major lab. The ratio measures machine runtime, not three human-equivalent days or a 3.1-fold productivity gain.
The company describes an automated research intern that can complete well-defined research tasks under human direction, including work estimated to take a skilled researcher days. Its measurements also show rising code activity and more experiments. People still set priorities, decide which results deserve further work and control whether a system is scaled, paused or deployed.
Agents Take More Execution Work
OpenAI defines an agent-workday as eight hours of machine runtime. At 3.1 agent-workdays per human workday, the reported organization-wide ratio corresponds to 24.8 hours of agent runtime for every eight aggregate hours of human labor. Agents can run concurrently, so the number describes how much machine time the organization used, not how much completed human work the systems replaced.
OpenAI’s research organization includes people who build research infrastructure, manage projects and support the research enterprise. The 3.1 figure is a mid-August 2026 snapshot rather than an average.
The systems used include OpenAI’s Codex coding agents, which are designed to complete delegated tasks for coding but also other types of work. Inside the research organization, OpenAI says classified agent activity widened from research and infrastructure coding into technical troubleshooting and monitoring runs. High-level planning remained a small part of that activity.
What OpenAI Calls a Research Intern
OpenAI defines its automated research intern as a system that completes well-defined research tasks under human direction, including some tasks estimated to take a skilled researcher several days. Under that definition, a longer task can count toward the milestone even when a person supplies the goal, intervenes during execution and decides whether the result merits further work.
OpenAI estimated task difficulty by how long a skilled person would need, then used an internal agentic classifier to identify success only when a ground-truth outcome could be found. Uncertain outcomes were excluded. The company says classifier-assigned success generally increased across several task-duration buckets from January through July 2026.
The human-time estimate and the runtime ratio answer different questions. A four-to-eight-hour task bucket describes expected difficulty for a skilled person; an agent-workday counts how long machines ran. METR’s methodology makes the same conceptual separation in its own, different task suite: estimated human completion time is a difficulty axis, not agent runtime or a measure of labor replaced.
Human involvement remained substantial in the harder bucket OpenAI discussed. Over the six months before the report, more than half of successful tasks estimated at four to eight human hours received at least one intervention. That denominator includes only tasks retained as successful after uncertain outcomes were filtered out. However, it does not reveal how many attempts failed, whether an intervention was a clarification or a rescue, or how much work the person contributed.
More Activity Is Not Proven Productivity
OpenAI says experiments per active experimenter reached their highest level since tracking began in January 2025, with the high arriving in August 2026. Code activity also increased, but the numeric code-output comparison lacks a stable unit and denominator.
Those changes show greater activity, while the cause remains mixed. Available compute grew substantially over the same period. More compute can support more experiments even if the effectiveness of each researcher or agent does not change, so the figures cannot isolate the contribution of coding agents.
The measures cover different objects. Runtime records machine use, while classified tokens indicate where agent activity occurred. Experiments per active experimenter use an unspecified definition of active. Classifier success covers only tasks with findable ground truth after uncertain outcomes were removed.
The Bottleneck Moves to Judgment
OpenAI’s figures describe a shift in researchers’ attention, not the removal of researchers. Agents can write and run code, troubleshoot infrastructure and monitor runs in parallel. OpenAI says people still set priorities and decide whether to scale, pause or deploy. They also judge which ideas and outputs deserve more resources.
Giving agents more execution work can therefore coexist with frequent intervention. Execution can expand while evaluation becomes the limiting step. A researcher who delegates several tasks may receive more candidate results, but someone still has to distinguish a useful result from a plausible-looking dead end and decide what enters the next experiment.
A 2028 Target Comes With Safety Conditions
OpenAI presents the intern milestone as a step toward its March 2028 target for an automated AI researcher. That target concerns a broader future role. The current measurements show bounded execution under direction, minimal participation in high-level planning and continued human judgment.
On August 7, OpenAI announced it tested Astra-class systems before the later release in higher-security environments. During the following week, one analyzed slice of reinforcement-learning workloads showed GPU allocation to Astra-class systems falling 59.2 percent, while other model classes absorbed about 85 percent of that decline. OpenAI interpreted the pattern as researchers substituting other models for Astra.
OpenAI chief scientist Jakub Pachocki links the 2028 ambition for an automated AI researcher to recursive self-improvement, in which AI systems take a larger role in improving later systems. He also now argues that labs should slow when alignment and monitoring cannot keep pace.
OpenAI’s report documents how they are already assigning more execution work to its agents. OpenAI’s 2028 target would move those systems from bounded tasks toward a broader researcher role, while people and institutions still decide which goals, results and risks justify continuing the work.


