Seven months ago, the amount of Anthropic’s own AI research that Claude was running by itself rounded down to nothing. Less than one percent. Today the company says that figure is 26 percent — not Claude helping, not Claude autocompleting, but Claude taking a high-level prompt and carrying most of a research task to the finish line while a human watches.

That is the headline number from the R&D Automation Index, a set of internal measurements Anthropic published this week alongside two companion metrics on agent oversight and compute allocation. It is the most specific public accounting any frontier lab has offered of how much of its own work the machines are now doing. It is also, and this matters enormously, a report card the company wrote about itself.

What “leads” actually means here

The word is doing a lot of load-bearing work, so it is worth pinning down. Anthropic borrowed its yardstick from Epoch AI, a research outfit that developed a six-rung Automation Level scale running from AL0 to AL5.

  • AL0 — no AI involvement whatsoever; a human does the work
  • AL3 — “collaborates”: the model handles large chunks under close human direction
  • AL4 — “leads”: the model completes most of a task end-to-end from a high-level prompt, with a human supervising
  • AL5 — fully autonomous, no human in the loop

The 26 percent refers to work rated AL4. The more striking number, arguably, sits one rung down: more than 90 percent of the measured R&D work now happens at AL3 or above. Claude is somewhere in the room for nearly everything Anthropic does to build Claude.

And the number Anthropic wants you to notice is the one at the top of the ladder. AL5 registered zero. “Claude is not operating fully autonomously for any measured subset of AI R&D work,” the company stated flatly. That line is not a throwaway. A lab that publishes a chart showing recursive self-improvement climbing steeply has a strong incentive to also publish the ceiling it has not yet hit.

Thirty thousand agents and a billion decisions

The scale disclosure is where the abstraction turns concrete. On its primary internal platform, Anthropic says roughly 30,000 agents were running research and engineering work simultaneously as of August. Those agents generated more than a billion individual decisions over the course of that month.

Every one of them passed through what the company calls online monitors — automated checks that evaluate an action before it executes. The block rate came in at 0.002 percent, or about one action in every 47,000. A second layer, the offline monitors, reviews transcripts after the fact, flagging roughly one to two in every thousand. That works out to around 100,000 transcripts reviewed weekly and about 50 escalated to a human each week.

Here is the honest read on that blocking rate: it supports two completely different stories. Either 30,000 agents are behaving themselves almost perfectly, or the monitors are calibrated loosely enough that only egregious actions trip them. Anthropic does not claim to know which. Neither does anyone else, because nobody outside the company has looked.

How the 26 percent got built

The methodology is more interesting than the number. Anthropic sampled 20 percent of staff across its model R&D departments during July, pulling from Slack messages and internal documentation to reconstruct what people were actually doing. Claude research agents combed through that material and produced a catalogue of roughly 15,000 granular tasks, which were then folded into a hierarchy of 542 nodes, 378 of them terminal.

A separate Claude instance acting as judge assigned an automation level to each one.

Read that again. Claude agents surveyed the work, Claude organized it, and Claude graded it.

Anthropic anticipated the objection and put numbers against it. The model’s ratings matched an employee’s exactly 59 percent of the time. Two employees rating the same task agreed with each other only 35 percent of the time. Model and human landed within one level of each other 97 percent of the time. On that evidence, the judge model is more consistent than the humans it is standing in for — which is reassuring right up until you remember that a judge sharing architecture with the systems it evaluates can be wrong in correlated, invisible ways.

The compute number nobody is quoting

The third index is the quietest and possibly the most consequential. During a one-week snapshot in mid-July, Anthropic measured how much of its R&D compute went toward safety work as opposed to capabilities.

The answer: about 6 percent across all AI R&D. Among the subset of research Claude itself is driving, it rose to 12 percent.

Anthropic describes both figures as deliberately conservative — ambiguous work was classified as capabilities by default — and cautions that a single week establishes nothing about direction. Fair enough. But it invites an uncomfortable follow-up. If the pitch is that automated research will help solve alignment faster than it creates alignment problems, six percent is a thin slice of the pie to be betting on.

The verification gap

Anthropic has done something genuinely unusual here. It adopted an external measurement scale it did not design, applied it to its own operations, published the results including the unflattering ones, and documented its own caveats in detail. That is more disclosure than the industry norm by a wide margin.

The readings themselves, though, are entirely self-generated. No independent party has audited the task catalogue, re-run the ratings, or checked the monitor logs. The company says it intends to embed third-party evaluators with access comparable to its internal risk teams. That is a plan, not a fact, and the distinction is the whole ballgame.

One methodological wrinkle deserves a flag. The task basket is frozen at a July baseline and will be rebuilt periodically, so the percentage gets re-versioned over time. Next quarter’s figure may not be comparable to this one.

What This Means

Strip away the framing on both sides and a trend line is hard to argue with. Under one percent to 26 percent in roughly six months is the kind of curve that makes forecasting look foolish. If it holds even loosely, the question of what a frontier lab’s research organization looks like in 2027 becomes genuinely open.

But the disclosure is not evidence that AI is building AI unsupervised. AL4 explicitly includes a human in the loop, and the zero at AL5 is real. What Anthropic has described is an organization where humans set direction and machines execute at enormous volume — a change in throughput and in who does the typing, not yet a change in who decides.

The part that should stick is structural rather than numerical. We now have a lab publishing quarterly-style metrics on its own automation, using a scale it borrowed rather than built, and naming in advance the outside verification it has not yet obtained. Whether that becomes a norm other labs adopt, or a one-off gesture that quietly stops getting updated when the numbers turn awkward, is the thing to watch. Self-reported transparency is better than none. It is not the same as being checked.