Automation and Agents
What Does Anthropic’s 26% AI-Led Research Figure Measure?
Anthropic says Claude led 26% of its measured AI R&D work in August 2026. The figure depends on task ratings and staff-time weights; here is how to read it.

On this page
- What does “AI leads” mean in this index?
- How did Anthropic turn its R&D work into a percentage?
- A worked example shows why the weights matter
- How should readers combine the three August findings?
- Where does judgment enter the 26% result?
- What would a rising index show over time?
- What would make another lab’s percentage comparable?
- How to read the next AI-led R&D percentage
Anthropic says Claude “led” 26% of its AI research and development work as of August 2026. That sounds like a measure of how much research Claude completed. It is more specific: Anthropic rated categories of its own R&D work by the role AI played, weighted those categories by estimated staff time, and found that 26% of the weighted work fell at a level called “AI leads.” Anthropic says no measured subset reached its fully autonomous level. The figures and method come from Anthropic’s measurement report.
The distinction matters before anyone compares this number with a future result. A rise could mean Claude handles more of familiar tasks. It could also involve revised task categories, weights, or judgments about where collaboration ends and leadership begins. Those possibilities call for different conclusions.
The useful reading is narrow: 26% is Anthropic’s prototype index result for its measured work in August 2026. It is neither a count of projects Claude finished nor a measure of fully autonomous research. To assess the next percentage, follow it through the work included, the rating scale, the weights, and the checks on those decisions.
Reading guide: Ask what work entered the index, which human decisions remain at “AI leads,” how much weight each category carries, and whether the method changed between measurements. A percentage without those details is difficult to compare.
View image detailWhat does “AI leads” mean in this index?
“AI leads” describes a way of carrying out a task while a person still supervises. Anthropic uses a six-level scale that it credits to Epoch AI. At AL3, AI collaborates: it performs substantial parts of a task under close human direction. At AL4, AI leads: it handles most of the task from a high-level request while a person supervises. AL5 is the fully autonomous level in Anthropic’s application of the scale. Epoch AI’s published rubric likewise distinguishes close human direction at level 3 from supervision and approval at level 4. Epoch’s publication explains the rubric; it does not verify Anthropic’s internal ratings.
Anthropic makes the boundary concrete with a broken nightly data pipeline. In its AL3 illustration, an engineer brings Claude the problem, helps resolve surprises during the investigation, reviews the fix, and decides whether to deploy it. At AL4, the engineer can hand Claude the failure alert. Claude investigates, repairs, tests, and documents the result without needing direction at every unexpected turn. The engineer still decides whether the change ships. In Anthropic’s AL5 illustration, Claude could notice the failure and deploy the repair without a person’s participation being required. These are Anthropic’s explanatory scenarios, not disclosed cases from which readers can recalculate the index.
The final approval in the AL4 example is consequential. Claude may lead the investigation and implementation while a person retains the deployment decision. Conversely, an AI system may produce a large volume of code while an engineer directs every important choice. Output volume alone does not establish that AI led the work.
Epoch’s authors describe their June 17, 2026 taxonomy and ratings as an initial, subjective attempt to track AI R&D automation. The two documents answer different questions: Epoch explains the level distinctions and the challenge of mapping research tasks; Anthropic reports how it applied a scale to its own R&D process. They are separate source documents, but they are not two independent measurements of Anthropic’s 26% figure.
View image detailHow did Anthropic turn its R&D work into a percentage?
Anthropic calls its prototype the Anthropic R&D Automation Index. The index combines three choices: which work to include, how to rate each kind of work, and how much weight each category receives. A change in any one can move the headline result. Anthropic describes those choices in enough detail to interpret the figure, though its published report does not include the underlying work records needed for an outside recalculation.
To build its work map, Anthropic says it randomly sampled 20% of staff in each department involved in its model R&D loop for each week of July 2026. A Claude research agent examined sampled staff members’ work records, including Slack and internal documents, and listed their tasks. Anthropic reports roughly 15,000 granular tasks from that process. Claude organized them into a hierarchy with 542 nodes, including 378 leaf categories. Anthropic then froze the tree as a basket for its measurements. The sampling and task-tree method appear in the report’s appendix.
A second stage assigned an automation level. Anthropic says a Claude agent researched how work in each category was performed, who performed it, which tools they used, and how much AI contributed. A separate Claude judge read that evidence and rated the category. For a given month, Anthropic says its research agents could see evidence from that month or earlier. That rule limits which records could inform an earlier month’s rating; it does not independently establish that each rating was correct.
The third stage gave categories different weights. In Anthropic’s July sample, each person contributed one unit per week, divided evenly among that person’s recorded tasks. Four tasks would receive a quarter-unit each; ten tasks would receive a tenth each. Anthropic added those allocations to form category weights. It calls person-time a crude proxy, intended to give more influence to work occupying more staff time.
That construction defines what the 26% can mean. It is a share of a staff-time-weighted map of work rated “AI leads.” It is not a count of completed research projects, a stopwatch measure of Claude’s labor, or a measure of the scientific value of Claude’s contributions. It also does not say Claude performed exactly 26% of the effort inside every category rated AL4. The rating applies to the category; its estimated weight determines how much that category contributes to the index.
A reader can use the method without accepting every judgment behind it. Anthropic has made the main components visible: staff sample, task tree, model-based ratings, and person-time weights. The public report does not provide the full set of category-level evidence and decisions needed to reproduce the result. That access limit should travel with the number whenever it is used as evidence.
View image detailA worked example shows why the weights matter
Imagine a hypothetical research group with 100 sampled person-week units across four categories. Infrastructure maintenance carries 50 units, evaluation design carries 30, experiment reporting carries 10, and tool maintenance carries 10. Suppose AI is rated AL4 for evaluation design and AL3 for the other three. The share at “AI leads” is 30%, because the AL4 category carries 30 of the 100 units. Counting categories instead would give one out of four. Neither calculation reports how many projects Claude completed.
Now suppose a rater moves the ten-unit experiment-reporting category from AL3 to AL4, with the work map and weights unchanged. The “AI leads” share rises from 30% to 40%. That move might reflect evidence that AI now handles surprises with less direction. It might instead reflect a different judgment about the same borderline work. A reader would need the reason for the rating change to distinguish those explanations.
A revised task map raises a separate question. If a broad category is split into narrower ones, its parts may receive different ratings. That could make the result more faithful to the work, but it would change how the new figure relates to the old one. A useful report would show which earlier categories became which later categories, and whether prior results were recalculated on the revised map.
This arithmetic is an illustration, not a reconstruction of Anthropic’s index. Rise has not inspected Anthropic’s category weights or repeated its ratings. The example shows why a percentage cannot be interpreted from its numerator alone. The denominator, weights, and classification decisions determine what moved.
Person-time is one way to weight a category, not a universal measure of importance. A brief task could affect a major research choice. A time-consuming task could have a modest effect on the final research outcome. Anthropic’s index answers a question framed around estimated staff time. If someone wants to know whether a task is worth automating in their own organization, Rise’s Work Worth Doing test considers judgment and failure consequences alongside repetition. That is a separate decision from interpreting Anthropic’s research index.
View image detailHow should readers combine the three August findings?
Anthropic reports that, as of August 2026, more than 90% of its measured AI R&D work was rated at or above “AI collaborates,” 26% was rated “AI leads,” and no measured subset was rated fully autonomous. These are Anthropic’s internal index results. They describe positions on one scale, so the first two percentages should not be added. Work rated AL4 is already within the broader share at or above AL3.
The more-than-90% result indicates widespread AI participation under Anthropic’s method, including categories where people remain closely engaged. The 26% identifies the narrower weighted share Anthropic rated AL4. The absence of a measured AL5 subset says the index did not rate any covered category as fully autonomous. Each finding narrows the interpretation of the others.
The unit is still a category of work, not an individual action. “No measured subset at AL5” does not tell readers whether every agent action received timely human approval. Nor does AL4 imply that Claude made every consequential choice in a category. The index describes the division of work at the task-category level; agent monitoring and action review require different evidence.
The observation date matters too. August 2026 is when the reported work was assessed. Anthropic’s research listing dates the report September 17, 2026, while the supplied capture of the report body does not display a publication date. The report body was captured for this draft on October 3, 2026. Those dates serve different jobs: observation, publication listing, and source capture. None turns the August estimate into a test of Claude performed by Rise.
View image detailWhere does judgment enter the 26% result?
The AL3–AL4 boundary is a judgment about how closely a person directs the work. Both levels can involve substantial AI output and a person who remains accountable. A task near the boundary might begin with a high-level request but require a consequential human choice midway through. Whether it qualifies as leadership depends on how raters apply “close direction” and “most of the task end to end.”
Anthropic says it asked staff responsible for relevant work areas to rate their areas without seeing the evidence or ratings produced by its model judges. It reports 59% exact agreement between model and human ratings, compared with 35% exact agreement between human raters. Model and human ratings were within one level of each other 97% of the time. The appendix reports these comparisons and acknowledges borderline cases.
The near-agreement figure can sound reassuring while leaving the headline threshold disputed. AL3 and AL4 are adjacent. If the model assigns AL4 to a heavily weighted category and a staff rater assigns AL3, their ratings are within one level, yet they disagree about whether the category belongs in the 26% share. The reviewed report does not state an uncertainty range showing how alternative plausible ratings would affect that share.
Agreement checks also address only part of the method. They do not independently establish that the sampled records cover all relevant R&D work, that the task tree groups unlike activities appropriately, or that person-time is the right weight for every question readers might ask. Anthropic identifies another concern: using its own models to judge its systems could allow a judge to share errors with the system it assesses.
The practical request for a later report is specific. Show which heavily weighted categories crossed the AL3–AL4 boundary, what evidence supported each move, and how sensitive the total is to disputed ratings. These are proposed reporting checks, not an audit Rise has carried out. They would help a reader tell whether the index moved because the work changed or because the interpretation of that work changed.
View image detailWhat would a rising index show over time?
A frozen task basket lets Anthropic compare ratings of work represented in that basket. If the same kind of task moves from AL3 to AL4 under consistent criteria, the index can reflect more AI leadership in that familiar work. Freezing categories protects one part of a trend: a newly named category cannot silently replace an older one without a visible method decision.
It creates a limit as well. Imagine a hypothetical team that automates routine experiment reports and uses its freed time to design a new kind of evaluation. A basket built around earlier work might register the reports becoming more automated while describing the new evaluation work poorly. A rising score on the old basket would then be an incomplete account of how the team’s work changed. This is an interpretation example, not an observation about Anthropic.
Anthropic reports a check on that concern. It built an alternate task tree from January 2026 data and compared tasks arriving from February through July 2026 with the January basket. It says it found no rise in “novel” tasks at its level of analysis, suggesting the structure of model R&D work was stable across that period. Anthropic plans to rebuild the basket periodically and re-version published numbers when appropriate. These statements appear in its appendix. The check applies to Anthropic’s chosen granularity and those months; it cannot establish that the basket will remain representative indefinitely.
Epoch’s authors make a related methodological point about task taxonomies: a list based on current human workflows may need revision as AI systems and research processes change. Their original proposal is context for why task maps need maintenance. It is not evidence that a particular missing category changed Anthropic’s result.
For any future time series, ask for a bridge between versions. Which categories stayed stable, which were split or added, and which earlier periods were recalculated? Were the person-time weights held fixed or refreshed? Did the rating criteria or judge change? A clearly labeled break can be more informative than a smooth chart that joins measurements made with different definitions.
View image detailWhat would make another lab’s percentage comparable?
Two labs could each say AI leads 26% of R&D while measuring different things. One might include infrastructure, evaluation, and product engineering; another might cover only model-training research. One might weight categories by staff time while another counts tasks. Their raters could also draw the AL3–AL4 boundary differently. Matching percentages would not settle any of those differences.
Anthropic names two obstacles to cross-lab comparison: no shared methodology and its use of its own models to evaluate its systems. It suggests verification by third parties or other developers’ models, with controls for competitively sensitive information. Those are proposals in Anthropic’s report, not completed checks of its August result against another lab. Epoch’s published scale offers a common vocabulary, but sharing level names alone does not make two work inventories or weighting schemes equivalent.
A meaningful comparison would specify the work covered, use compatible rating criteria and units, align observation periods, and disclose how task categories and weights were constructed. It would also give an independent reviewer enough access to inspect samples of underlying work, challenge category assignments, repeat selected ratings, and check the calculation. A review limited to the final chart could confirm its presentation without verifying how the chart was built.
Anthropic says it plans to embed outside evaluators with access comparable to its internal risk assessment teams. The report reviewed for this draft does not establish that this arrangement has begun or that an evaluator has reproduced the automation index. When a later publication cites independent review, ask which measurement was checked, what records the reviewer saw, and which parts remained outside its access.
View image detailHow to read the next AI-led R&D percentage
Begin with the claim’s scope and period. Identify the organization, the activities included, and the month or other window from which the ratings come. Anthropic’s figure concerns its measured AI R&D work in August 2026. It cannot be carried over to another organization or a later month without a new measurement.
Then identify the unit. Is “work” a share of weighted task categories, a count of tasks, staff hours, or compute? Anthropic’s result uses category ratings and estimated person-time weights. Another publisher could use the word “leads” while calculating something different.
Read the human role beside the level. Can AI handle surprises from a high-level request? When must a person direct, correct, or approve it? Anthropic’s AL4 pipeline illustration leaves deployment with an engineer. That detail is essential to the meaning of leadership in its report.
Finally, examine the change record and verification. Which categories or weights moved? How were disputed ratings resolved? Did an outside reviewer inspect source records and rerun parts of the calculation, or only review a summary? A future percentage is most useful when it explains what changed in the work separately from what changed in the instrument.
Anthropic’s 26% offers a structured baseline for discussing Claude’s role in its measured R&D work. Epoch’s published rubric helps explain the levels, while Anthropic’s report supplies the internal estimate. Neither document shows that Claude fully autonomously builds its successor or establishes a comparable rate at another lab. The next meaningful update would show the level distribution, category and weight changes, remaining human decisions, and completed verification alongside the headline figure.
For action-level decisions, use the AI agent oversight checks before expanding authority. Task-level AI leadership does not establish that an agent’s consequential actions are adequately monitored or approved.
Checked for this article



