Monday, Sep 7, 2026

Top stories

Research acceleration: The view inside OpenAI

OpenAI published its first detailed internal metrics on how coding agents are changing research work: the median researcher now runs agents daily, burning over $600/day in inference at API prices (the 90th-percentile user tops $7,000/day), and as of mid-August the research org logs 3.1 "agent-workdays" of effort for every workday of human labor — up from below parity before June. Task success rates on longer, harder tasks are rising, but over half of successful 4-8 hour tasks still needed human intervention, and delegated work is shifting toward higher-level, longer-horizon tasks. The post also discloses the compute cost of safety: after the recent Hugging Face incident, OpenAI paused and then restricted RL training on its latest (Astra-class) models, cutting that class's compute allocation by double digits before other model classes largely absorbed the difference.

Sources & depth

An Alien Mind

OpenAI chief scientist Jakub Pachocki argues the current scaling trajectory could sustain into recursive self-improvement within a few years, and says plainly that no lab has yet solved alignment and monitoring well enough to keep scaling at maximum speed much longer. Notably, he discloses that chain-of-thought monitoring — OpenAI's primary technique for checking whether models are reasoning toward misaligned goals — is becoming less reliable as models get smarter and blend reasoning with tool use and multi-agent interaction. His prescription isn't just internal caution: he calls for external, mandated safety bars enforced by third-party auditors, government agencies, or international bodies, not voluntary lab-by-lab restraint.

Sources & depth
Also today