New Transparency Framework for Frontier Labs
Anthropic has released a comprehensive set of internal metrics designed to provide external visibility into the pace of AI development. The Anthropic report, published in September 2026, focuses on three core pillars: the extent to which AI models are automating their own development, the efficacy of oversight for autonomous agents, and the allocation of compute resources toward safety versus capability.
As frontier AI models become increasingly powerful, the industry faces mounting pressure to demonstrate that development remains under human control. Anthropic’s initiative seeks to establish a replicable framework that other frontier labs could adopt, potentially serving as a foundation for future regulatory oversight or industry-wide pacing coordination.
Measuring AI-Led Research and Development
One of the most significant metrics introduced is the “Anthropic R&D Automation Index.” By utilizing an automation rating scale (AL) developed by Epoch AI, the company evaluated the role of its Claude models in internal research. As of August 2026, the data indicates that while Claude does not operate fully autonomously, it currently “leads” 26% of Anthropic’s AI R&D work—meaning the model completes tasks end-to-end with human supervision. Furthermore, over 90% of internal R&D work involves AI collaboration, where models handle significant chunks of the project under close human direction.
The company acknowledges that this reliance on AI to build future AI models creates a “recursive self-improvement” challenge. To mitigate risks, Anthropic plans to embed independent, third-party evaluators within its labs, granting them access to internal processes and data to verify safety practices.
Agent Oversight and Compute Allocation
The report also details the oversight regime for the approximately 30,000 agents operating on Anthropic’s internal platforms. Monitoring is split into two categories: online, which blocks dangerous actions in real-time, and offline, which reviews activity retrospectively. Data from August 2026 shows that 100% of agent actions pass through these monitors. During the study period, only 0.002% of over one billion decisions were blocked by online monitors, suggesting that individual agent misbehavior remains rare but requires rigorous, scalable oversight protocols.
Finally, the company highlights compute allocation as a verifiable input for tracking development. By categorizing workloads, Anthropic aims to show how much compute is dedicated to safety-critical work, such as model alignment and interpretability, versus raw capability scaling. The firm suggests that compute transparency could become a key lever for governments to monitor or pace the development of frontier models industry-wide.

