Anthropic Says Claude Leads 26% of Its AI R&D: Inside the New Automation Index
Anthropic published a new set of internal AI-development measurements on September 17, 2026, reporting that Claude leads 26% of its measured AI research and development work as of August. More than 90% of the measured work is at or above the company's “AI collaborates” threshold, while none of the measured work reached its fully autonomous level.
The same disclosure says approximately 30,000 agents were doing research and engineering work concurrently on Anthropic's most-used internal agent platform in August. Anthropic reports that 100% of actions on that platform pass through an online monitor before execution and are also ingested by an offline monitoring system afterward.
Anthropic is proposing three measurements for frontier AI development: the share of R&D performed by AI, oversight of internal AI-agent actions, and the allocation of AI R&D compute. The company presents them as prototype transparency metrics that other frontier developers could publish using comparable methodologies.
The headline measurements
| Metric | Anthropic-reported result | Measurement scope |
|---|---|---|
| AI “leads” R&D work | 26% | Measured Anthropic AI R&D, August 2026 |
| AI “collaborates” or above | More than 90% | Measured Anthropic AI R&D, August 2026 |
| Fully autonomous AI R&D | 0% | Measured Anthropic AI R&D |
| Concurrent internal agents | About 30,000 | Most-used internal agent platform, August 2026 |
| Online-monitor coverage | 100% of agent actions | Same internal platform |
| Offline-monitor ingestion | 100% of agent actions | Same internal platform |
| AI R&D compute allocated to safety | About 6% | Examined week, July 13–20, 2026 |
| AI-driven AI R&D compute allocated to safety | About 12% | Same examined week |
These figures are company measurements, not an industry benchmark. Anthropic publishes methodological detail that makes the categories inspectable, while cross-lab comparison would require common definitions and independent verification.
How the R&D Automation Index works
Anthropic calls its first measurement the R&D Automation Index. It uses an automation scale developed by Epoch AI ranging from AL0, with no AI involvement, to AL5, where AI operates fully autonomously without a human in the loop.
At AL3, labelled “collaborates,” AI performs large portions of work under close human direction. At AL4, labelled “leads,” AI can complete most of a task end-to-end from a high-level prompt while a human supervises.
To construct the underlying task map, Anthropic says it sampled 20% of staff in each department involved in model R&D for each week in July. A Claude research agent reviewed work records including Slack and internal documentation and produced roughly 15,000 granular tasks. Claude then organized those tasks into a hierarchy containing 542 nodes, including 378 leaf categories. Anthropic says it freezes that tree so subsequent measurements use the same basket of work.
The resulting August snapshot places 26% of measured work at AL4 and more than 90% at AL3 or higher. The AL5 share is zero.
Agent oversight at 30,000-agent concurrency
The second measurement focuses on how Anthropic supervises internal agents. Its most-used research and engineering platform had roughly 30,000 agents running concurrently in August, according to the company.
Anthropic separates monitoring into online and offline systems. Online monitors evaluate actions before execution and can block or redirect them. Offline monitors analyze activity after execution for patterns that may develop over longer periods.
The company reports 100% coverage for both paths on the measured platform. Automated online review occurs before execution, with humans reviewing blocked actions within one week. The offline system flags roughly 100,000 transcripts per week for additional classifier review, with approximately 50 highest-priority cases per week escalated to humans.
Coverage measures whether activity passes through a monitor; it does not by itself measure the monitor's detection accuracy. Anthropic identifies monitoring effectiveness and independent verification as areas needed for stronger external interpretation of these metrics.
Measuring safety compute
Anthropic's third proposed metric tracks how AI R&D compute is allocated. For the week of July 13–20, the company classified about 6% of AI R&D compute as safety-related and about 12% of AI-driven AI R&D compute as safety-related.
Anthropic describes these as conservative estimates. Work contributing equally to capabilities and safety was excluded from the safety total, and safeguards-classifier compute was outside this measurement.
Compute share is also an incomplete proxy for safety effort because many safety tasks are researcher-intensive without consuming training-scale compute. Its practical value is as a consistently defined quantity that can be tracked over time if developers publish their category definitions.
Why these measurements matter
Frontier labs increasingly use their own models to accelerate model research, evaluation, infrastructure and engineering. That creates a measurable feedback loop between current AI capability and the speed at which successor systems can be developed.
Anthropic's 26% AL4 result gives that trend a concrete internal measure. The 30,000-agent concurrency figure adds an operational dimension: oversight systems increasingly have to supervise machine activity at a scale where manual review of every action is infeasible.
The disclosure also exposes the main obstacle to comparison. Anthropic currently uses its own models in parts of the measurement process, and labs can classify work differently. A common methodology plus third-party access to underlying systems and records would make cross-company measurements substantially more useful.
Anthropic says it plans to embed independent third-party evaluators from multiple organizations with access comparable to internal risk-assessment teams. Those evaluators are intended to verify safety practices, report incidents and monitor metrics such as the ones introduced here.
Bottom line
Anthropic's new measurements provide one of the more detailed public snapshots of AI-assisted frontier-model development: Claude leads 26% of measured R&D work, participates at the collaboration threshold or above in more than 90%, and operates across an internal platform running about 30,000 research and engineering agents concurrently.
The durable value of the framework will depend on repeated measurements, stable definitions and external verification. If other frontier developers publish compatible data, the automation index, agent-oversight metrics and compute allocation could become useful indicators of how quickly AI is being incorporated into the process of building subsequent AI systems.
Sources
- Anthropic Institute — Measurements for understanding the pace of AI development inside frontier labs: https://www.anthropic.com/institute/measuring-pace-of-ai-development
- Epoch AI — automation-level methodology referenced by Anthropic: https://epochai.substack.com/
- Anthropic Responsible Scaling Policy: https://www.anthropic.com/responsible-scaling-policy
- The Rundown AI — independent coverage of Anthropic's measurements: https://www.therundown.ai/news/anthropic-claude-ai-research-transparency