Claude now leads a quarter of Anthropic’s R&D

Artificial Intelligence 5 min 18 Sep 2026
Bar chart rising to 26 percent Claude leads Anthropic AI R and D with agent grid
Anthropic's R&D Automation Index: Claude leads 26% of AI R&D, collaborates on 90%+, zero full autonomy. ~30k agents, 0.002% blocked.

Anthropic’s new R&D Automation Index puts Claude R&D automation on the public record: Claude now leads about 26% of the lab’s AI research and development work, up from under 1% in March. More than 90% of that work sits at “AI collaborates” or higher on Epoch AI’s scale, and zero measured subsets are fully autonomous. Roughly 30,000 research agents run on the main internal platform, with online monitors blocking about 0.002% of more than a billion August decisions.

That is the hard number the pacing debate has been missing. Yesterday’s desk covered OpenAI’s misalignment disclosure and the DeepMind Institute essays; today the lead is Anthropic measuring how far Claude already drives the next Claude. I am reading it next to Amodei’s call to coordinate on pace, and next to the recursive self-improvement scare word labs keep dancing around.

Let’s dig in.

Bar chart rising to 26 percent Claude leads Anthropic AI R and D with agent grid

Image: From near-zero “leads” share to 26% by August 2026: Source: oguzhan.co editorial composite

📊 Claude R&D automation hits 26% “leads”

The primary source is Anthropic’s Institute post, Measurements for understanding the pace of AI development inside frontier labs. They catalogued AI R&D work into a frozen basket (about 15,000 sampled tasks organized into 542 categories), rated each category on Epoch AI’s Automation Level scale from AL0 to AL5, and weighted by person-time.

AL3 means AI “collaborates”: large chunks under close human direction. AL4 means AI “leads”: most of the task end-to-end from a high-level prompt, human still supervising. AL5 is full autonomy with no human in the loop. As of August 2026 Anthropic reports Claude leads 26% of AI R&D, collaboration-or-higher sits above 90%, and AL5 is still empty on every measured subset.

Business Insider and Reuters amplify the same snapshot: the “leads” share was below 1% in March. Six months is a brutal slope if you are trying to argue humans still own the research loop.

My caveat, which Anthropic also owns: the judge is Claude rating Claude-assisted workflows. Staff spot-checks agreed with the model about as often as humans agreed with each other, and ratings landed within one level 97% of the time. Good enough for a public pace thermometer. Not a court exhibit yet. That is why the third-party evaluator embed plan matters.

Epoch Automation Levels AL0 AL3 AL4 AL5 ladder for Anthropic R and D

Image: Collaborates above 90%, leads at 26%, autonomous still empty: Source: oguzhan.co

🛡️ Thirty thousand agents, 0.002% blocked

The second measurement is oversight, not vibes. Anthropic says about 30,000 agents do research and engineering on the most-used internal platform at any one time. Online monitors see 100% of those actions before execution. Offline monitors ingest everything after the fact.

Over August they analyzed more than a billion agent decisions. Online monitors blocked 0.002%, roughly one in 47,000. Offline monitors flag one to two transcripts per thousand; about 50 high-priority cases escalate to humans each week. Coverage is complete on that platform. The open question is whether “blocked” and “flagged” keep meaning anything when agent count and decision volume climb another order of magnitude.

Compute share is the third dial. In a mid-July week, about 6% of AI R&D compute went to safety work, and about 12% of the compute inside AI-driven AI R&D was safety-tagged. Anthropic calls those numbers conservative. They also note safety work is often researcher-time heavy and compute light, so the percentage is a comparator, not a virtue score.

Agent oversight snapshot showing 0.002 percent block rate and safety compute share

Image: Oversight coverage is high; the block rate is tiny; safety compute share is mid-single digits: Source: oguzhan.co

⏱️ Pacing, RSI, and what “leads” actually buys

Anthropic frames the index as a way to see how close the world is to recursive self-improvement: a model fully autonomously building its successor. They also say the numbers would shift if labs coordinated on pacing the frontier, the line Amodei has been pushing in public.

This sits inside the week’s cluster. Wednesday’s deep dive asked whether an AI standards body is a pact or a cartel. Yesterday’s digest tracked OpenAI’s misalignment disclosure and the DeepMind Institute essays. Today’s index is the missing instrument panel: not “are we worried,” but “what share of the factory is already AI-led.”

I want three receipts before I treat this as industry-wide fact. Other labs publishing the same AL breakdown. Third parties with internal access verifying the ratings. And a shared definition of safety compute that survives marketing. Anthropic says it plans to embed independent evaluators with access comparable to internal risk teams. That is the test. Until then, Claude R&D automation is a loud, useful, single-lab meter.

🧪 Desk note: Bend wants proofs, not vibes

One lighter HN signal while the labs argue about pace: Bend pitches a language that blocks AI coding mistakes via proof, running on CPU and GPU. I am not rewriting my toolchain tonight. I am filing it next to the Anthropic story because the industry has two parallel answers to “AI writes the next AI.” One answer is more agents and better monitors. The other is making wrong code harder to ship. Both can be true. Neither replaces a public automation index.

📡 Signals

Watch list into the weekend: whether OpenAI or Google publish comparable AL4 shares; whether Anthropic’s third-party embed timeline gets dates; whether the 26% “leads” number becomes the new citation in every pacing hearing; and whether Bend-style verified tooling shows up in any lab’s internal agent scaffold. A 26% lead share with zero AL5 is still human-supervised acceleration. It is also the steepest public slope we have seen this year.

That was Friday’s digest. Numbers over slogans. 🙋‍♂️

Oğuzhan Koçaklı

I have worked professionally in marketing, gaming and blockchain since 2015. I have helped create and carry out marketing strategies for many major brands. These days I work on mobile games and blockchain integration for games. AI has been my hobby for many years.

All posts

Leave a Reply

Your email address will not be published. Required fields are marked *