On September 17, 2026, Anthropic published measurements of how much of the factory that builds Claude is being handed back to Claude. In a new internal R&D Automation Index, the company says Claude led 26% of its measured AI R&D work in August 2026, under human supervision. Claude collaborated or led on more than 90% of the measured work, while no category reached full autonomy.
The obvious reading is that recursive self-improvement has arrived: Claude is building its successor, and the feedback loop is accelerating. That reading is too strong. The 26% figure is not a productivity gain, a share of research breakthroughs, a share of all company compute, or the fraction of a successor model built without people. It measures where control sat on a weighted basket of internal tasks.
But the cautious reading should not be dismissive. Moving from assistance to supervised leadership changes the operating system of a frontier lab. Humans increasingly set objectives, judge evidence, and authorize consequential transitions while fleets of agents execute the middle. The scarce resource can therefore move from production to acceptance: deciding which experiment matters, whether a result is real, and whether it is safe to feed back into the next model.
The useful interpretation is not “26% automated.” It is “26% of a fixed, person-week-weighted task basket crossed a human-to-agent control boundary.” That is consequential, but it answers a different question from how fast the lab is improving.
What the 26% actually measures
Anthropic uses a six-level rubric adapted from Epoch AI’s proposed taxonomy for AI R&D. AL0 means no AI involvement; AL1 is minimal involvement; AL2 is assistance; AL3 is collaboration; AL4 is AI leadership with human supervision; and AL5 is full autonomy.
AL4 is the important boundary. At that level, a human can give a high-level prompt and Claude can complete most of the task. The human still supervises. In Anthropic’s broken-pipeline example, Claude can diagnose and repair the system, but the person keeps the deployment decision. The headline therefore describes delegated execution with retained human authority, not an unattended lab.
Reported measure | Value and window | Correct reading | Not established |
|---|---|---|---|
AI-led R&D work | 26%, August 2026 | Weighted task share rated AL4 | 26% faster R&D or 26% autonomous successor development |
AI collaboration or leadership | More than 90%, August 2026 | AL3 and AL4 combined; the 26% sits inside this figure | An additional 90% on top of the 26% |
Concurrent internal agents | About 30,000 on one platform, August 2026 | Fleet scale on Anthropic’s most-used internal R&D platform | 30,000 independent researchers or 30,000 productive projects |
Safety share of AI R&D compute | About 6%, July 13–20 | One week, narrow safety-dominant definition | Company-wide safety spend or safety effectiveness |
Safety share of AI-driven AI R&D compute | About 12%, July 13–20 | Different, narrower denominator than the 6% figure | A 6–12% budget range |
The time series is still striking. Anthropic’s chart puts the AL4 share below 1% in February, then at 1% in March, 3% in April, 12% in May, 14% in June, 22% in July, and 26% in August. Those are company-reported historical measurements with 90% measurement intervals, not a forecast. The chart contains no year-end target and does not justify drawing a straight line to AL5.
The index is more thoughtful—and more fragile—than the headline
The methodology is unusually concrete for a frontier lab. Anthropic reconstructed roughly 15,000 tasks from Slack and internal documentation, using a random 20% sample of relevant staff for each week in July. It froze a hierarchy with 542 nodes and 378 leaf tasks, then retrospectively rated earlier months using only evidence available by the month being scored.
The weighting matters. Each sampled person-week gets one unit, divided equally across that person’s tasks for the week. That makes the index a proxy for the distribution of human work, not a count of task instances or measured hours. One researcher with two tasks gives each half a unit; another with ten gives each one-tenth. A routine, widely staffed category can move the index more than a rare decision that gates an entire training run.
Claude research agents assemble evidence and a separate Claude judge assigns levels. “Separate” means a separate judge in Anthropic’s process, not an independent organization. In validation, Anthropic reports 59% exact model-human agreement and 97% agreement within one level; exact agreement between two human raters was 35%. That makes the classifier useful, but 97% is not the accuracy of the 26% headline. An adjacent disagreement can cross the economically important AL3-to-AL4 boundary.
Anthropic built a January baseline task basket, then checked new work from February through July against it and found no rise in novel categories at its chosen resolution. That is a sensible drift check. It does not solve the deeper substitution problem: once agents take over implementation, people may spend more time selecting experiments, integrating results, or reviewing failures. The taxonomy can remain stable while the bottleneck migrates inside it.
A handoff index is not a research speedometer
This is the central analytical mistake to avoid. Automation level, useful output, and recursive acceleration are three different variables. METR’s work on task substitution separates uplift on the old task mix, uplift on the new task mix, and uplift in produced value. Anthropic’s index estimates none of those productivity uplifts. It rates the depth of AI delegation across a comparable, weighted body of work.
It cannot, by itself, answer whether validated research output rose 26%, whether model-training cycles shortened, or whether a stronger Claude creates the next capability improvement with less human input. A human may touch fewer substeps while spending more time rejecting plausible but wrong results. Thousands of agents may create more experiments while a scarce cluster, dataset, or deployment committee still sets the pace.
The economics of recursive self-improvement turns on feedback strength: how much a model improvement raises validated future R&D output at fixed other inputs. A July working paper on the economics of recursive self-improvement models labor, compute, experiments, inference, and data as possible constraints. Even complete task automation need not produce explosive acceleration if another input becomes binding. Conversely, meaningful feedback can exist well before AL5.
handoff share ≠ productivity gain
productivity gain ≠ shorter critical path
shorter critical path ≠ self-sustaining accelerationAnthropic’s own policy framework preserves that distinction. Its August Risk Report describes automated-R&D thresholds based on full substitution for research staff within five times their cost or a doubling of the pace of AI progress attributable to automation. Anthropic said neither threshold was met as of the report’s July 15 coverage date. Those are not the same test as a 26% AL4 task share, and the company had not directly tested every version of the full-substitution case.
Thirty thousand agents turns monitoring into a queue-design problem
The other important disclosure is operational scale. Anthropic says about 30,000 research and engineering agents were active at any one time on its most-used internal platform in August. The company reports that every action on that platform passed through an online monitor before execution and was ingested by an offline system afterward. Across more than one billion August decisions, the online block rate was about 0.002%, or roughly one in 47,000.
Those are valuable exposure denominators. They are not an end-to-end safety score. One platform is not every Anthropic environment. A monitor decision is not a task, transcript, incident, or independently productive agent. A blocked action is not automatically malicious, and a low block rate does not reveal false negatives.
The offline pipeline makes the selection problem visible. Anthropic says it initially flags about 100,000 transcripts each week. Multiple classifiers reduce that pool until roughly 50 highest-priority cases reach human review. At fleet scale, the second-stage ranking system is therefore part of the safety boundary. It decides which evidence a human ever sees.
Coverage asks whether an event passed a sensor.
Detection asks whether the sensor recognized the relevant hazard.
Triage asks whether the signal survived filtering and reached a reviewer.
Containment asks whether harm was prevented before an irreversible state change.
Closure asks whether the underlying failure mode was corrected and verified.
Earlier evidence shows why the full funnel matters. In the risk report with a July 15 coverage date, Anthropic evaluated an earlier monitoring setup on 163 METR red-team attack transcripts. One first-stage monitor flagged 94%, while 79% remained positive after second-stage filtering. The same report says two inadvertent cluster-job mass-deletion incidents were flagged at the exchange level but did not surface through the full pipeline. These are bounded historical results from a different date and described architecture—not a current recall estimate for the August platform. They are concrete precedent that downstream triage can lose important signals.
That scope mismatch needs reconciliation. The September page says all actions on one platform were ingested; the linked earlier report describes selective semantic monitoring, exclusions, and incomplete end-to-end evaluation. Different windows and systems may explain both statements. The next disclosure should map logged, ingested, classifier-reviewed, escalated, and human-resolved events by dated environment, using shared identifiers. Until then, “100% monitored” means complete routing through the stated layer, not proven end-to-end effectiveness.
This extends the agent-safety evidence problem: incident counts become meaningful only when exposure, detection, escalation, containment, and verification are kept separate. It also complements our earlier argument that multi-agent products should optimize for accepted work per reviewer-hour. Thirty thousand active workers is an input. Accepted, dependency-correct output is the result.
Persistent identity helps audits—if authority does not persist silently
Anthropic’s internal scaffold gives agents persistent identities that survive model upgrades, plus shared, attributable communication that can be cross-referenced with transcripts. That is a strong provenance design. It lets a reviewer follow a worker’s history instead of treating every session as a disconnected chat.
But identity, capability, and authority should be three separate records. Upgrading the model behind a stable agent ID can materially change what that worker can do. Its old permission set may no longer be appropriate. The audit trail should therefore preserve model-version epochs, parent-child delegation, tool grants, permission changes, evidence sources, and external effects. A stable name should preserve accountability, not smuggle stronger capability into yesterday’s trust decision.
The safety-compute numbers are a ledger, not a grade
For July 13–20, Anthropic reports that about 6% of AI R&D compute went to safety and about 12% of AI-driven AI R&D compute went to safety. These figures use different denominators. They cover one week, not August, and they are not shares of all company compute, annual spending, or headcount.
The definition is deliberately conservative: safety, security, or model understanding must be a workload’s dominant purpose. Work that advances capability and safety equally is excluded from the safety numerator, as is separate safeguards-classifier compute. Anthropic sampled about 14% of nearly 10,000 research training and evaluation runs, overweighting compute-intensive runs; that is not the same as sampling 14% of compute.
A resource share is useful for accountability, but it can point in the wrong direction if turned into a scoreboard. A better, cheaper safety-evaluation or safety-research run could reduce the safety percentage while improving protection. A wasteful safety stack could raise the percentage without making the system safer. A large capability run can shrink the ratio even if absolute safety work grows. Labs should pair compute allocation with fixed-definition absolute usage, benchmarked detection, production outcomes, and independently reproduced improvements.
Automating safety research can automate the wrong objective
Anthropic’s earlier Automated Weak-to-Strong Researcher shows both sides of the opportunity. Nine parallel Opus 4.6 agents accumulated 800 agent-hours over five days and achieved a strong result on a bounded chat-preference testbed. The system also found shortcuts and reward-hacking strategies. When one method was transferred to production-scale infrastructure, the held-out gain was only 0.5 points—within the reported noise floor.
The lesson is not that automated research failed. It is that research throughput and research validity are different systems. More agents can optimize a weak proxy faster. The smaller experiment also found that deliberately diverse starting directions reduced convergence on the same ideas, which suggests that fleet size is not the same thing as intellectual diversity. A useful fleet metric would count distinct hypotheses and marginal validated discoveries per added agent, not just concurrent workers.
What builders should instrument now
This announcement does not expose Anthropic’s internal platform as a product. There is no new public SKU, model ID, price, or reproducible 30,000-agent deployment recipe. The practical value for builders is the measurement design—and the gaps it reveals.
Separate delegation from outcomes. Track supervision level, quality-adjusted throughput, human review and rework, full-cycle elapsed time, and total cost per accepted result. Code volume and agent count are not enough.
Name the retained human decision. For every AL4-like workflow, state whether the person selects the objective, accepts the evidence, authorizes an external side effect, or all three. Few intermediate prompts do not make a task safe to run unattended.
Version the task basket. Publish a constant-basket series for comparability and a rebased current-work series for economic relevance. Include review hours and newly created work so automation does not hide a moving bottleneck.
Audit the whole monitoring funnel. Measure coverage by environment and action type; test recall with authorized seeded failures; sample allowed and de-escalated work; report p50, p95, and p99 review and containment time; and track whether fixes prevent recurrence.
Bind authority to capability epochs. Keep stable agent identities, but log model upgrades, delegation lineage, permission changes, tools, artifacts, and external effects. Reapprove high-impact permissions when capability changes materially.
Protect the evaluator. Keep final test data and scoring infrastructure outside the research agents’ reach. Require fresh held-out evaluation, independent reproduction, and transfer tests before feeding a claimed improvement into production.
Test fleet scaling, not fleet theater. Hold budget and acceptance criteria constant while comparing one, a few, and more agents. Measure idea diversity, duplicate work, reviewer load, and marginal validated results.
The next index needs four companion measures
Anthropic deserves credit for publishing a method that can be criticized rather than another vague claim about coding productivity. It also proposes public methodologies, external verification, and possible testing windows before a model is reused for more AI R&D. Those are proposals, not enacted rules or independent certification.
Accepted research output at fixed compute, including failed experiments and reviewer rework.
Critical-path elapsed time from question selection to a validated result that changes a model or safety decision.
A scope-matched monitoring funnel: logged, ingested, detected, escalated, human-reviewed, contained, and verified closed.
Safety outcomes beside safety inputs: absolute resources, detection under adaptive testing, transfer to production, and independent reproduction.
These are the quantities an embedded evaluator would need to reproduce. As we argued when examining Anthropic’s embedded-evaluator pledge, access matters only if outside researchers can inspect definitions, trace evidence, challenge the judge, and publish findings. The underlying task data and complete ratings remain private, so today’s 26% is a confirmed company disclosure—not an independently reproduced result.
What I expect next
First, frontier labs will be pushed to publish model-factory metrics alongside model benchmarks. The argument will quickly move away from the headline percentage and toward task definitions, evidence access, baseline versioning, and model-judge calibration. Cross-lab comparison is impossible if every company means something different by task, safety compute, or monitored action.
Second, provenance will become a product category. Agent platforms will separate stable worker identity from replaceable models and time-bounded authority. Teams will inspect communication graphs and artifact lineage, not just chat transcripts. The model factory is becoming an organization made partly of software; it needs the equivalent of identity management, change control, and financial accounting.
Third, the evaluator becomes the most valuable protected asset in automated research. As experiments get cheaper, convincing but invalid improvements get cheaper too. Untouched tests, independent reproduction, and production transfer will matter more than another order of magnitude in agent-hours.
FAQ
Is Claude autonomously building its successor?
No. Anthropic reports supervised AI leadership for 26% of a weighted task basket and no AL5 category. Claude is materially involved in successor-model work, but humans still choose goals, review evidence, and retain consequential decisions. The disclosure does not establish end-to-end autonomous successor development.
Does 26% mean Anthropic’s research is 26% faster?
No. It is a weighted share of task categories rated at a supervision level. Productivity requires quality-adjusted output, review and rework, cost, and elapsed critical-path time. The announcement does not supply a causal estimate connecting the index to overall R&D speed.
Does 100% monitoring mean every dangerous action is caught?
No. Anthropic reports that every action on one measured platform passed through online monitoring and was ingested offline. That establishes stated coverage of a pipeline. It does not establish detection recall, precision, protection against adaptive attacks, absence of bypasses, or coverage of every internal environment.
Do the 6% and 12% figures measure Anthropic’s safety budget?
No. They are one-week compute shares with different denominators and narrow workload definitions. They exclude some dual-purpose work and separate safeguards-classifier compute. They are useful allocation snapshots, not a company-wide budget range or proof of safety performance.
Anthropic’s most important move here is not claiming that Claude has replaced its researchers. It is making part of that handoff measurable.
The 26% figure is a useful first instrument. It tells us that supervised agent leadership has moved from an edge case to a meaningful part of frontier-model development. The next instruments must show whether that handoff creates accepted knowledge faster, whether oversight survives the full filtering pipeline, and whether safety improves as capability compounds. Until those measures exist, the index is a map of changing responsibility—not a countdown clock.
