Lifestyle

Anthropic Publishes Three Frontier AI Metrics: 26% Is “Leads,” Not “Fully Autonomous”

On September 17, 2026, Anthropic published three metrics meant to let outsiders track R&D progress inside a frontier AI lab: how automated its research is, how AI agents are overseen, and how compute is allocated to safety. Fact-checked on September 18, 2026, this article explains why 26% “leads” is not full automation, why the 0.002% block rate carries its own caveat, and why these are still self-measured numbers with limited scope, as external evaluators are still being set up.

About 14 min read

Original illustration: three gauge dials with needles at different readings, and a geometric question-mark outline above them, showing these readings still await outside scrutiny
Image: Mokaair (© Mokaair)

On September 17, 2026, Anthropic published three metrics on its official website meant to let outsiders see research and development progress inside a frontier AI lab: how much AI R&D work AI itself completes, how the actions of AI agents are overseen, and how compute resources are allocated. Anthropic ties this to an earlier piece by CEO Dario Amodei calling for pacing frontier AI, writing that it would expect these numbers to shift if labs actually coordinated on slowing down.

This article was fact-checked on September 18, 2026, against three pages on Anthropic's official website: the post announcing these three metrics, the newsroom overview page, and a policy page explaining the policy context. This site has not run any tests of its own; all three metrics are currently numbers that Anthropic measures and publishes itself, and this article goes through the limits and the parts that remain unverified, as Anthropic itself has flagged them.

What Anthropic Published, and Why

The summary Anthropic's official newsroom page wrote for this post is that today, the world can't see what's going on inside AI labs, and Anthropic is proposing new metrics to give the public visibility into frontier AI development. The post itself ties this to Amodei's call, but only links out to his piece without restating what it argues; Anthropic writes that these proposed measurements and policies together are a starting point for monitoring the pace of AI development from outside the labs.

The post itself does not print a publication date — as fact-checked through September 18, 2026, the page has no byline date, no “published on” text, or corresponding metadata field. September 17 is the date Anthropic's official newsroom page attaches to this piece in its featured section, and this article treats that as the event date; the original text prints no time of day, so this article does not state an exact publication time. Anthropic calls this set of metrics a prototype, writes that it hopes to demonstrate transparency this way and will keep doing so, but does not commit to a fixed publishing schedule.

Three Metrics: R&D Automation, Agent Oversight, and Safety Compute

The first is the “Anthropic R&D Automation Index,” scored using Epoch AI's Automation Level (AL) scale, running from AL0 to AL5 (fully autonomous, with no human in the loop). At AL3, “collaborates,” AI can do large chunks of work under close human direction; at AL4, “leads,” AI can complete most of a task end-to-end from a high-level prompt while a human still supervises. As of August 2026, Anthropic writes that Claude is not operating fully autonomously for any measured subset of AI R&D work; it “leads” 26% of the work, and the share at or above “collaborates” totals more than 90%.

The second is agent oversight, where Anthropic picked three angles: coverage, review latency, and escalation rate, corresponding respectively to how much of an action passes through monitoring, how long elapses between an action and its review, and how much of an action gets blocked or flagged for review. As of August 2026, roughly 30,000 agents were doing research and engineering work at any one time on Anthropic's most-used internal platform, and these figures cover only that one platform. 100% of these actions pass through online monitoring first, usually within seconds, and afterward 100% also go through offline monitoring; the offline monitor flags roughly 100,000 transcripts a week, most of which are resolved by classifiers, with the highest-priority roughly 50 a week escalated to human review.

Anthropic writes that after analyzing more than a billion decisions in August 2026, 0.002% of them (about 1 in 47,000) were blocked by the online monitor. At the start of that section, Anthropic first cautions that in its monitoring data to date, individual agents rarely misbehave, but when there are millions or billions of agents operating in the economy, even rare events can happen regularly. The third metric is a compute snapshot: Anthropic examined how it used all of its compute over the week of July 13 to 20, 2026, finding that about 6% of AI R&D compute was allocated to safety work and about 12% of “AI-driven AI R&D” compute was allocated to safety work; the two figures have different denominators, and this article does not add or compare them.

Compiled from the three metrics Anthropic published on September 17, 2026; fact-checked September 18, 2026. All figures in the table are Anthropic's own measurements.
MetricPeriodOfficial FigureLimit Anthropic States
R&D Automation IndexAugust 2026“Leads” at 26%; “collaborates” or above more than 90%Claude scoring Claude, not an outside audit
Agent Oversight: CoverageAugust 2026100% of actions pass online monitoring beforehand, offline monitoring afterwardCovers only one internal platform, the most-used one
Agent Oversight: Block RateAugust 20260.002% blocked out of more than 1 billion decisions (about 1 in 47,000)Anthropic says rare events can happen regularly at economy scale
Safety Compute ShareJuly 13–20, 2026About 6% of AI R&D compute, about 12% of AI-driven R&D compute, allocated to safety workAnthropic calls it a deliberately conservative single-week snapshot

Why These Numbers Are Not Yet Verified Numbers

Anthropic itself lists two obstacles standing in the way of cross-lab comparison. First, there is no shared measurement methodology; second, Anthropic uses its own Claude to score its own systems: Claude research agents first gather the evidence, and then a separate Claude serves as the judge; when rating a given month, the agents doing the rating can only see evidence from that month or earlier. In other words, Anthropic itself writes that this could mean the model acting as judge makes the same kinds of errors as the model it is checking.

The oversight numbers have boundaries too: Anthropic writes that it plans to embed independent third-party evaluators from multiple organizations, giving them access comparable to what its internal risk assessment teams have, and in the oversight section writes that it is now setting up external evaluators — neither sentence amounts to something an outside party has already checked. The text mentions that METR has, in the past, independently red-teamed Anthropic's offline monitoring platform, but that is past tense, and what was tested was the monitoring platform, not the percentages this article publishes. Based on what this fact-check read on the day it was done, this article found no outside party that has verified the numbers above.

Anthropic also writes that its method for the compute share is conservative: the example it gives is that if a token contributed as much to advancing capabilities as it did to safety, it was not counted on the safety side; the two figures also exclude a separate, similarly sizable pool of compute used for safety-protection classifiers. Anthropic notes that safety research tends to rely more on researchers' time than on compute compared with frontier training, so the compute share is only an imperfect proxy, and its value lies in comparing like with like across developers and over time, not in the absolute size of these two percentages by themselves.

Four-panel diagram: who does the R&D, how agents are overseen, who compute is allocated to, and who verifies it
A summary of the three metrics Anthropic published on September 17, 2026: R&D automation, agent oversight, and the share of compute spent on safety. Fact-checked September 18, 2026. · Image: Mokaair (© Mokaair)

The Two Numbers Most Likely to Be Misread

The first place readers are likely to misread is “leads.” 26% is often taken to mean “a quarter of the work is done by AI on its own, with no human managing it,” but Anthropic's own definition of AL4 “leads” is that AI can complete most of a task end-to-end from a high-level prompt while a human still supervises; AL5, “fully autonomous, with no human in the loop,” is described in a footnote as “a level we have not yet reached.” A closer reading is “AI takes on a task and does most of the steps itself, but someone is still watching and can step in at any time,” not “AI finishes the job behind closed doors, with no one knowing what happened.”

The second place readers are likely to misread is the 0.002% block rate. Taken alone, it is easy to read as “agents almost never misbehave, so the system is safe”; but in that same section, Anthropic writes that when there are millions or billions of agents operating in the economy, even if any single event is unlikely, rare events can happen regularly. And 0.002% is only the rate the online monitor blocked within one month on one internal platform; the offline monitor separately flags roughly one to two in a thousand transcripts for further review, and the original text lists the two in separate cells of a table without combining them.

What This Article Leaves Out, and How It Connects to Pacing the Frontier

Based on what was fact-checked through September 18, 2026, this post does not name any country or market, and it does not mention product changes or any feature change users would see — this is an article about how Anthropic measures its own internal R&D process, not a product announcement, and it will not change how Taiwan-based users use Claude today. The article publishes only Anthropic's own numbers; it does not give other companies' automation share, oversight coverage, or safety-compute share. Anthropic's own text says “any frontier developer could” publish similar metrics, not that other companies already have.

In an appendix, the post writes out its own definition of “safety work”: work whose dominant purpose is making AI systems safer, more understandable, or more secure; Anthropic also states that this is just one reasonable choice among many, that another developer or regulator might draw the line elsewhere, and that it believes all sides would benefit from agreeing on a shared definition ahead of time. A separate Anthropic policy page describes its own Advanced AI Framework, covering transparency and independent evaluation up through government authority to block or deter dangerous deployments when necessary; that page states that the framework is written primarily with the US federal government in mind and is a proposal, not a rule already in effect.

This post and the piece Amodei published earlier, “We Must Pace the Frontier,” belong to the same thread: the post only writes that these numbers echo Amodei's call for coordinated pacing of the frontier and links out to it, without restating its content; the three-step plan and “step one is embedding external evaluators” come from another roundup on this site, not from the pages cited in this article. The three metrics published here are Anthropic's attempt to turn “whether pacing is actually happening” into numbers that can be laid out and examined; to learn about those three steps, see the second link at the end of this article.

Frequently asked questions

Does 26% mean a quarter of R&D work has already been handed entirely to AI?

No. In Anthropic's own definition, 26% corresponds to Automation Level AL4, “leads,” meaning AI can complete most of a task end-to-end from a high-level prompt, but a human is still supervising; AL5, fully autonomous with no human in the loop, is described in a footnote as “a level we have not yet reached,” and the text also states that as of August 2026, no measured subset of AI R&D work is yet fully autonomous.

Does the 0.002% block rate mean this oversight system is safe?

Anthropic does not claim that. This figure is the rate the online monitor blocked out of more than a billion decisions in August 2026, and at the start of that same section Anthropic writes that in its monitoring data to date, individual agents rarely misbehave, but when there are millions or billions of agents operating in the economy, even if any single event is unlikely, rare events can happen regularly. The offline monitor separately flags roughly one to two in a thousand transcripts, which is a different figure that the original text does not combine with the block rate.

Have these three metrics been verified by an outside party?

Based on what this article found as of its September 18, 2026 fact-check, no outside party has verified these numbers themselves. Anthropic writes that it plans to embed independent third-party evaluators with access comparable to its internal risk assessment teams, and elsewhere writes that it is now setting up external evaluators. The text mentions that METR has, in the past, independently red-teamed Anthropic's offline monitoring platform, but that is past tense and what was tested was the monitoring platform, not these percentages.

Does this have anything to do with Claude users in Taiwan?

Based on what this article found, the post does not name any country or market, and it does not mention product changes or any feature change users would see. This is an article about how Anthropic measures its own internal R&D process, not a product announcement, and it will not change how Taiwan-based users use Claude today.

Will these metrics be published on a fixed schedule going forward?

Anthropic has not committed to a fixed publishing schedule. The post says it hopes to demonstrate transparency this way and will keep doing so, and also that any frontier developer could publish this kind of metric on a regular basis using a public method; as of the page this article fact-checked, there is no date given for the next release, nor any promise to keep using the same method going forward. In an appendix, Anthropic also writes that it plans to periodically rebuild that basket of tasks and re-version its published automation numbers as appropriate.

How does this piece relate to Amodei's earlier call to pace frontier AI?

This Anthropic post writes that these numbers echo Amodei's call for coordinated pacing of the frontier and links out to his piece, without restating its content. The three-step plan and “step one is embedding external evaluators” come from another roundup on this site, not from the pages cited in this article. The three metrics published here are Anthropic's attempt to turn “whether pacing is actually happening” into numbers that can be laid out and examined; to learn about those three steps, readers can continue with the second link at the end of this article.

  • Lifestyle

    Amodei Urges Pacing Frontier AI: What “We Must Pace the Frontier” Proposes and Its Limits

    Anthropic CEO Dario Amodei published “We Must Pace the Frontier” on September 12, 2026, arguing for slowing the advancement of AI capabilities so that risk prevention has time to keep up, with third-party evaluators confirming this. This article covers how “pacing” is defined and what it excludes, the three-step plan, how Altman, Hassabis, and Musk each responded, and what this advocacy piece actually means for everyday users.

  • Lifestyle

    OpenAI Publishes a Misalignment Reporting Framework: All 6 Reports Come From Training

    OpenAI's own page prints a publication date of September 16, 2026, U.S. time, which is already September 17 in Taipei time; the same day it also released 6 case reports. Checked against the framework announcement, the alignment.openai.com overview page, and two of the case reports, this article explains what stage and which models the 6 reports are marked with, what steps a disclosure goes through, and what each of the two published rates actually measures.

  • Lifestyle

    Anthropic's September Threat Report: Seven Kinds of AI Misuse and Targeted API Keys

    Anthropic published a threat intelligence report on September 10, 2026, covering seven categories of AI misuse it disrupted between December 2025 and August 2026. This article explains what the report says and what it does not, and how ordinary people can guard against AI-enabled scams, protect their accounts and API keys, limit AI agent permissions, and report suspicious use.

  • Lifestyle

    NVIDIA launches DGX Spark 64GB: on sale October 23 from $4,999, two units can be linked into 128GB

    On October 2, 2026, NVIDIA announced a more affordable 64GB memory version of its DGX Spark personal AI computer, available from October 23 through six makers including Acer and ASUS. It is aimed mainly at developers and researchers who want to run AI models on their own machines. Below we summarize the specs NVIDIA published, its claims about linking two units, and what it means for general readers.

Latest travel guides

Sources

Lifestyle