When AI Starts Building the Next AI

AI is already doing a growing share of the work behind the next generation of AI. As development accelerates, the question is whether human oversight and external scrutiny can keep pace.

Sep 24, 2026 - 07:32
Uppdaterad: 16 timmar sedan
0 4
A human analyst views abstract AI evaluation displays in a control room.

What long seemed like a thought experiment about AI has become a practical question inside the companies building the most advanced models: how much of the work on the next generation of AI is already being done by AI systems?

Anthropic has tried to put numbers on that question. The company says Claude “led” 26 percent of the AI research and development work included in its internal measurement in August 2026. More than 90 percent of the measured work was at a level where AI at least collaborates with people. No measured part of the work was at the level of full autonomy.

The figures say something important about change inside AI labs. But they do not mean that Claude is building the next Claude on its own, nor are they an external certification of how far automation has advanced. As AI does more research work, who measures the progress - and can oversight grow as quickly?

Key points

Anthropic reports that Claude leads 26 percent of measured AI R&D and that more than 90 percent is at least at the collaborative AL3 level. No measured part is at the fully autonomous AL5 level. The measurements are largely internal, not an external certification. OpenAI reports a parallel increase in AI use in research, but with a different methodology.

A change inside AI labs

In its account of how the pace of AI development can be measured, Anthropic presents a prototype “R&D Automation Index.” It maps tasks in model development, assesses how automated they are and combines the results.

This is different from what a language model can do in a single test. It concerns the ongoing work behind a new model: code, infrastructure, experiments, evaluations and other work needed to build, train and improve AI.

Anthropic says Claude “leads” 26 percent of this measured work, and that more than 90 percent is at or above “AI collaborates.” Its chart shows that the share led by Claude was below one percent in February 2026. The crucial word is not “26 percent,” but “leads.”

What does it mean for AI to “lead”?

Anthropic uses an automation scale developed by Epoch AI. It runs from AL0, no AI involvement, to AL5, where AI performs the work fully autonomously without a human in the loop.

AL3 means AI collaborates with a person. At AL4, the level Anthropic calls “leads,” a person can give a high-level assignment and let the system carry the work forward, handle problems and document results. A person is still meant to review the work and make important decisions, including whether a change goes into production.

AL5 is different: the system would identify a problem, define the task, carry out and test the solution, and put it into production without human involvement. Anthropic explicitly says none of the measured AI R&D work is at that level.

That distinction matters. AI taking a broad assignment and carrying much of the work through to completion is not the same as it deciding which research problems are worth solving, setting its own goals or independently developing a successor. In a separate article on recursive self-improvement, Anthropic says it is not there today and that this development is not inevitable.

30,000 agents at once

Anthropic says around 30,000 agents were simultaneously carrying out research and engineering work on its most-used internal agent platform in August. The figure applies to that platform, not automatically to all AI use within Anthropic.

The company describes online monitors that check actions before execution and can block or redirect dangerous actions, and offline monitors that review activity afterward for patterns that may signal problems over time.

Anthropic says 100 percent of these agents’ actions pass through the online monitor before execution and 100 percent are ingested by the offline monitor afterward. That means coverage of the flow, not that the system detects everything it should.

More than one billion research and engineering-agent decisions were analysed in August, according to Anthropic. 0.002 percent - roughly one in 47,000 - were blocked by the online monitor. The offline monitor flags around 100,000 transcripts per week; about 50 of the highest-priority cases go on to human review.

Those figures may reflect agents that rarely attempt anything dangerous, effective restrictions, or the permissions and environments in which the agents work. Without a measure of missed problems - recall and false negatives - a low blocking rate is not, on its own, a safety rating.

Analysts monitor a large number of parallel AI agent flows in a technical environment.
As thousands of AI agents work in parallel, the question is not only what they can do, but whether monitoring and human oversight can scale with the activity.

Who measures the pace?

Anthropic is unusually open that the measurement is internal. It built its task map from sources including Slack and internal documentation. For each week in July 2026 it sampled 20 percent of staff in the relevant parts of the organisation, producing roughly 15,000 detailed tasks in a tree with 542 nodes.

Anthropic also freezes the task basket so that automation can be compared over time. A rising index therefore shows how work in the July baseline is being automated, but does not automatically capture new kinds of work that may emerge as people move on to different tasks. Anthropic says a comparison with a January baseline did not show an increase in novel task types through July, but also says the basket will need to be rebuilt and versioned over time.

Claude organised the tasks and examined how work was done; a separate Claude model assessed automation. Anthropic also asked responsible employees to make blind assessments. The company says model and human agreed exactly in 59 percent of cases, while two human assessments agreed exactly in 35 percent; within one step, model and human agreed in 97 percent.

That is a disclosed methodological check, not a replacement for external review. When the same model family is used both in the operation and in parts of the assessment, the judge may share some errors with what is being assessed.

Epoch AI developed the scale and a broader taxonomy of AI R&D. It has not independently established that Claude performs 26 percent of Anthropic’s research. Comparisons between labs need a common method and some form of outside scrutiny.

When oversight also has to scale

As more agents can work in parallel, review becomes just as concrete a question as capability. How quickly can a person understand a flag? How do you catch a risk pattern invisible in a single decision but emerging from many agents together?

Microsoft Research has highlighted a related problem: an agent system can reach a seemingly correct final result even though it breached permission boundaries or leaked information along the way. Its study found that the risk surface grew in multi-agent environments and that safety depended heavily on the system allocating tools, resources and messages.

That does not prove anything about Anthropic’s monitors. It does explain why counting blocked actions alone is too narrow: as operations scale, questions are needed about what monitors see, how they are tested against evasion and when a person has time to intervene.

External review - what METR actually saw

In March, research organisation METR wrote that one staff member had spent three weeks red-teaming part of Anthropic’s internal systems for agent monitoring and safety, with substantial access to relevant systems and information.

The exercise found several new vulnerabilities, some later fixed, according to METR. It also said that none seriously undermined the central claims in Anthropic’s then-current sabotage-risk report.

External reviewers can find concrete weaknesses, and a red-team exercise can offer more than an internal dashboard. But this report is historical and limited; it is not a certification of Anthropic’s monitoring systems in September 2026.

Embedded evaluators - independent on whose terms?

On 18 September, Anthropic announced a collaboration with Accenture, led by its specialist business Faculty, on embedded evaluation. According to Anthropic’s description, the work will cover model evaluation and red teaming, alignment assessments and testing of safeguards.

External evaluators could receive access comparable to an employee’s, allowing them to follow models during training, see development and deployment decisions and speak directly with employees. That is potentially more visibility than a reviewer gets once a model is finished.

But “embedded” is not automatically the same as independent. Anthropic says standards for access, reporting and funding do not yet exist. Anthropic will fund Accenture’s evaluation work directly, while discussing trials with independent funding for METR and other nonprofit evaluators. Separately, Anthropic and Accenture each expect to invest at least $1 billion in building capacity in this area over five years.

Who chooses the reviewer, what access is guaranteed, what may be published and who decides when a concern should be made public? The model could bring better visibility, but it could also show how difficult it is to build an oversight function close to the organisation it scrutinises while remaining independent of its interests.

An external evaluator reviews test results and system flows from an AI development environment.
External evaluation can provide greater insight into how advanced AI systems are developed, while also raising questions about access, funding and who decides what can be made public.

OpenAI is moving in the same direction - but with different metrics

In a report from 6 September, OpenAI writes that its research organisation used 3.1 agent-days for every human workday in mid-August, against a standard eight-hour day. It is a measure of aggregate agent runtime, not an estimate of equivalent human productivity.

OpenAI says it has reached its goal of an “automated research intern”: a system that, under human direction, can perform clearly defined research tasks that might take a skilled researcher several days. It describes a goal of a more automated AI researcher by March 2028.

But it stresses that people still set priorities, choose which ideas and results to pursue, and decide whether to scale, pause or use the systems. In its internal data, more than half of successful tasks estimated to take people four to eight hours required at least one human intervention.

OpenAI’s 3.1 figure cannot be compared with Anthropic’s 26 percent as a percentage. The methods differ. Together, they show AI becoming a larger layer of research work while human direction remains visible.

Other labs, and the missing comparison

Outside the two companies, the public picture is more fragmented. Meta has published research on AIRA₂, an AI research-agent system that reaches or exceeds human results on some benchmark tasks. But the Meta study is a benchmark result, not an account of how much of Meta’s internal AI R&D is done by AI.

In NextNet’s research through 23 September 2026, we did not identify comparable public internal statistics from Google DeepMind, xAI, Nvidia, Mistral or Hugging Face that would allow a serious comparison. If automation is to become a measure of the field’s pace, it is not enough for each company to describe its own operation with its own figures.

Slow down or accelerate?

In “We Must Pace the Frontier”, Anthropic CEO Dario Amodei argues that capability development should leave more time for safety work and external verification, rather than require a total halt to model training and technical progress.

OpenAI has argued for shared ways of measuring progress toward more automated AI research, to preserve human control and to slow down or stop when safeguards are insufficient. In its policy article, it also says fully autonomous recursive self-improvement is not happening today.

In a Reuters report on 18 September, Mistral argued that pacing and regulatory models could favour already dominant US actors. Hugging Face CEO Clément Delangue opposed pacing while supporting embedded evaluators, while AI professor Kristian Kersting warned of regulatory capture. Reuters also describes the broader European debate in terms of digital sovereignty and the need to build domestic AI capacity.

That does not remove the need for scrutiny. More open metrics, stronger external evaluation and better incident reporting can be sought without accepting that a small number of companies should define the pace of development for everyone else.

When AI starts building the next AI

The most striking element in Anthropic’s and OpenAI’s accounts is not proof that people have left the research loop. It is how much of that loop can already be parallelised: debugging, coding, experiments, technical support and parts of evaluation.

That can make development faster and add capacity for safety work. But the same scale makes companies’ own accounts of oversight harder to accept uncritically. When thousands of agents work at once, the quality of monitors, testing and human oversight becomes as important as the agents’ capabilities.

Fully autonomous recursive self-improvement has not been shown to be a present-day reality. Intelligence-explosion scenarios remain scenarios, not established facts. Yet AI is visibly taking a growing role in building the next generation of AI.

The difficult question is not only how much research AI can do. It is who can verify the claims, what systems miss and whether accountability and visibility can scale as quickly as the development work itself.


💬 What do you think?

If AI does an increasing share of the work involved in developing the next AI, how much of the oversight should sit outside the companies themselves?

Share your thoughts in the comments.

Vad är din reaktion?

Gilla Gilla 0
Ogilla Ogilla 0
Kärlek Kärlek 0
Rolig Rolig 0
Wow Wow 0
Ledsen Ledsen 0
Arg Arg 0

Kommentarer (0)

User
Staffan Carlsson

Hej, jag heter Staffan Carlsson

Jag är grundare och ansvarig utgivare för NextNet.se – en svensk nyhetsplattform med fokus på artificiell intelligens, teknik, cybersäkerhet och digital innovation.

Varför NextNet?

Jag startade NextNet med målet att skapa en modern och lättillgänglig nyhetssajt där teknik och AI står i centrum. Den tekniska utvecklingen går snabbare än någonsin, och jag tror att det är viktigare än någonsin att kunna förstå vad som händer – utan att behöva vara expert.

Genom NextNet vill jag lyfta fram nyheter, analyser och trender som hjälper läsare att navigera i en allt mer digital värld.

Mitt teknikintresse

Teknik har varit en stor del av mitt liv under många år. Jag fascineras av hur innovation, internet och artificiell intelligens förändrar sättet vi arbetar, kommunicerar och bygger framtidens samhälle.

Utöver arbetet med NextNet ägnar jag mycket tid åt webbplattformar, servermiljöer, AI-lösningar och digitala projekt där nyfikenhet och lärande alltid står i centrum.

Min vision

Jag vill att NextNet ska vara en trovärdig och inspirerande källa för alla som vill följa utvecklingen inom AI, teknik och digitalisering – oavsett tidigare kunskapsnivå.

AI-drivna nyheter för en digital värld.