# Cambridge report examines whether AI-led research could accelerate AI progress

> A Cambridge working paper examines whether AI-led research could accelerate AI progress. Anthropic's 26% figure shows supervised work, while the proposed feedback loop remains conditional.

By BIG CHANGE Editorial

Published: 2026-09-28T23:34:31.964Z
Updated: 2026-09-28T23:34:31.964Z
Canonical: https://bigchange.ai/blog/casp-report-ai-automating-ai-research

![Conceptual charcoal illustration, seen from above, of one researcher comparing two sheets at a workbench while a desktop monitor faces away.](https://bigchange.ai/api/media/file/casp-research-review-hero-v2.png)
AI-generated conceptual illustration by BIG CHANGE.

Anthropic says Claude now leads 26% of the AI research and development work it measures inside the company. That means a system can complete most of a task from a high-level prompt while a person supervises. Anthropic also says Claude does no measured R&D work fully autonomously. The distinction matters in a new [working paper from the Cambridge Programme on AI Science & Policy](https://casp.ac/__l5e/assets-v1/5efd4b41-deb5-4513-a0a3-b4f82d2b79ea/intelligence-explosion.pdf), which asks whether AI doing more of its own R&D could eventually make AI progress accelerate sharply.

The paper, dated September 2026, brings together 22 authors from universities, AI companies and civil society. It argues that a rapid feedback loop is possible: capable systems help produce better systems, which then do more research. Its central conclusion is a warning about a plausible path and the need to measure it. The paper does not report that such an acceleration has begun.

## The big change

- **What changed:** A cross-institution working paper has set out the evidence for AI-led R&D, the conditions under which it might accelerate AI progress, and specific measurements it asks governments to obtain from frontier labs.
- **Why it matters:** More research tasks are being delegated to AI systems inside the companies developing them. Public measures of that delegation are still partial, making it hard to tell how much faster the entire research process is becoming.
- **What to watch:** Independently checked measures of how much work AI completes, whether its contributions improve later models, and whether compute, experiments and human review slow the feedback loop.

## The evidence is real, but the measures answer different questions

[Anthropic's September measurement report](https://www.anthropic.com/institute/measuring-pace-of-ai-development) says the share of its R&D work at the “AI leads” level reached 26% in August 2026, up from less than 1% in February. Its rating scale still places a human supervisor in that category. Anthropic says the index is a prototype, its own models help rate the work, and there is no common method for comparing such numbers across companies. Its statement that no measured work is fully autonomous deserves to sit alongside the 26% figure.

[OpenAI's September account](https://openai.com/index/research-acceleration-view-inside-openai/) describes researchers using more coding agents, contributing code faster and running more experiments. OpenAI says experiment volume is correlated with greater agent use, while available compute has also grown. People still choose research priorities and decide which results to pursue, it says. More code and experiments can aid research without establishing a matching increase in the rate at which more capable models are produced.

Independent evaluations help measure one part of the process. [METR's Time Horizon 1.1 update](https://metr.org/blog/2026-1-29-time-horizon-1-1/) finds that models can complete longer research and software tasks than earlier systems in its test suite. METR also reports wide uncertainty for long tasks, many of whose human completion times were estimated. A task benchmark measures capability on its selected tasks; it does not measure an AI lab's total research productivity. In a separate randomized study of 16 experienced open-source developers working on 246 tasks, [METR found that access to early-2025 AI tools increased completion time by 19%](https://metr.org/Early_2025_AI_Experienced_OS_Devs_Study-paper.pdf). That older software setting cannot stand in for a frontier lab in September 2026, but it shows why tool use and task benchmarks alone cannot settle the productivity question.

## A possible feedback loop has several brakes

The CASP paper concentrates on a software-driven loop. AI systems would expand the effective research workforce by doing work in parallel. Successful improvements would make the next systems better researchers, allowing more improvements. The authors say preliminary evidence is consistent with this mechanism and judge substantial R&D automation likely within a few years. Their stronger possibility of an “intelligence explosion,” compressing years of progress into months or less, depends on the loop outrunning its constraints. It is a conditional scenario, not a measured forecast of what will happen.

The paper names four important constraints: diminishing returns from additional research effort, limited compute and data, tasks that remain hard to automate, and processes such as long training runs that cannot be sped up at will. It says evidence about some of these is mixed or indirect, and that direct evidence on the strength of time-intensive bottlenecks is lacking. Its example of progress accelerating tenfold within about 1.5 years is a model calculation after full R&D automation, conditional on estimated returns to research effort persisting and no new bottleneck appearing. It is not an observed acceleration or an unconditional date.

Other work sharpens the uncertainty. In a [2025 study of four AI labs](https://arxiv.org/abs/2507.23181), Parker Whitfill and Cheryl Wu found different answers depending on how they modeled research compute. In one specification, cognitive labor and compute looked substitutable; in a model that accounted for the rising demands of frontier experiments, they looked complementary. The authors caution that the data and model are limited. [Toby Ord's August 2026 analysis](https://arxiv.org/abs/2608.14426) identifies the time needed to complete each improvement cycle as another critical parameter. These analyses do not rule out rapid acceleration. They show that the speed of the loop cannot be read directly from the number of agents or the amount of code they write.

## What the authors want measured

The Cambridge authors propose three policy priorities: obtain visibility into AI R&D automation, develop ways to steer and constrain a possible acceleration, and prepare for its potential effects. The most immediate and testable of these is measurement. They suggest reporting the share of research contributions produced by AI, the pace of algorithmic improvement, how labs allocate spending among people and compute, the research decisions AI systems influence, and incidents involving internal agents. They also propose independent auditing, with stronger requirements for systems at the frontier of AI R&D capability.

Some later proposals, including requirements before further deployment, ways to pause specified R&D workloads, and international verification, would involve substantial design and trade-offs. The paper itself warns that poorly designed powers could favor one company and delay beneficial work. Its account of possible benefits and severe harms rests on the earlier conditional acceleration. Policymakers can examine the near-term reporting case without assuming the most extreme scenario has occurred.

The useful question after this report is narrower than whether AI will suddenly become superintelligent. It is whether labs can show, in comparable terms, which parts of research AI now leads, which improvements survive human evaluation, and how fast those improvements feed into the next generation of systems. The public evidence is not yet sufficient to answer that question across companies.

## Sources & further reading

- [Cambridge Programme on AI Science & Policy, full working paper (September 2026)](https://casp.ac/__l5e/assets-v1/5efd4b41-deb5-4513-a0a3-b4f82d2b79ea/intelligence-explosion.pdf): The authors' argument, evidence review, model assumptions, limits and policy proposals. This is labeled a working paper; its public materials do not establish external peer review.
- [CASP executive summary](https://casp.ac/__l5e/assets-v1/026ee81c-773c-4811-9c97-f7cf367ee7b0/intelligence-explosion-summary.pdf): A shorter statement of the authors' conditional scenario and proposed responses. It was written by a subset of the authors.
- [Anthropic, measurements of AI-led R&D](https://www.anthropic.com/institute/measuring-pace-of-ai-development): The primary source for the 26% “AI leads” figure, its supervision level and its methodological limits.
- [OpenAI, research acceleration inside OpenAI](https://openai.com/index/research-acceleration-view-inside-openai/): Company-reported agent use, experiment volume, human direction and caveats about measuring research progress.
- [METR, Time Horizon 1.1](https://metr.org/blog/2026-1-29-time-horizon-1-1/): Independent task evaluation and uncertainty about how benchmark horizons are estimated.
- [Whitfill and Wu, compute bottlenecks](https://arxiv.org/abs/2507.23181), and [Ord, dynamics of intelligence explosions](https://arxiv.org/abs/2608.14426): Separate analyses of conditions that could slow or alter a recursive research loop.

## Sources

- [CASP, What if automating AI R&D triggers an intelligence explosion?](https://casp.ac/__l5e/assets-v1/5efd4b41-deb5-4513-a0a3-b4f82d2b79ea/intelligence-explosion.pdf) — Primary 14-page Frontier AI Working Paper Series No. 2/2026. Read in full. Mechanism, mixed evidence, assumptions and policy proposals.
- [CASP executive summary](https://casp.ac/__l5e/assets-v1/026ee81c-773c-4811-9c97-f7cf367ee7b0/intelligence-explosion-summary.pdf) — Two-page summary by a subset of the authors. Checked against full paper.
- [Anthropic, Measurements for understanding the pace of AI development inside frontier labs](https://www.anthropic.com/institute/measuring-pace-of-ai-development) — Primary company measurement of 26% AI leads, zero measured full autonomy and methodological caveats.
- [OpenAI, Research acceleration: The view inside OpenAI](https://openai.com/index/research-acceleration-view-inside-openai/) — Primary company measures of agent use and research activity; distinguishes correlation from net model progress.
- [METR, Time Horizon 1.1](https://metr.org/blog/2026-1-29-time-horizon-1-1/) — Independent task horizon evaluation with uncertainty and task-composition limits.
- [METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity](https://metr.org/Early_2025_AI_Experienced_OS_Devs_Study-paper.pdf) — Randomized study in a different software setting; bounded counterexample to assuming productivity from tool adoption.
- [Whitfill and Wu, Will Compute Bottlenecks Prevent an Intelligence Explosion?](https://arxiv.org/abs/2507.23181) — Model-based analysis with divergent compute-labor substitution estimates and explicit data limitations.
- [Toby Ord, The Dynamics of Intelligence Explosions](https://arxiv.org/abs/2608.14426) — Theoretical analysis of feedback generation time; not an empirical forecast.
