Automated AI R&D (recursive self-improvement)
Use AI agents to do AI research itself; OpenAI says it has an 'automated research intern' and targets an automated AI researcher by March 2028.
If agents can design experiments, write code and choose ideas, research speed compounds and capability growth accelerates. OpenAI reports 3.1 agent-workdays per human workday inside its research org; others pursue it through evolutionary code search, AI-scientist loops and autoresearch.
Where it stands. OpenAI claims 3.1 agent-workdays per human workday; METR judges AI R&D at Anthropic unlikely to be dramatically accelerated so far. OpenAI's next declared milestone is an automated AI researcher by March 2028.
Evidence
For
- OpenAI (2026-09-06): its research org logged 3.1 agent-workdays per human workday by mid-August; the median researcher spends $600+/day on inference at API prices (company-reported).
- AlphaEvolve (2025-05-14) sped a Gemini kernel 23%, cutting Gemini training time 1%, and a Borg heuristic recovers ~0.7% of Google's fleet compute.
- Self-reported researcher uplift: ~4x in Anthropic's Mythos Preview system card and ~2x in METR's Becker et al. (2026), both flagged as unreliable (METR, 2026-07-21).
- Karpathy's autoresearch (2026-03-06): an agent edits training code and runs 5-minute experiments overnight, an open template for agent-run research orgs.
- Sakana's AI Scientist is published in Nature (2026-03-26) after an AI-written paper passed ICLR-workshop review (2025-03-12).
- A preliminary METR report puts Anthropic's AI-driven acceleration near 1.5x, with a 30% chance of 2x (cited 2026-09-22).
Against
- METR's Opus 5.5 evaluation (2026-09-22) judged AI R&D at Anthropic unlikely to have been dramatically accelerated and full automation of AI R&D unlikely for this model; part of its evidence came from a separate team that could not share details, plus a source it could not disclose.
- OpenAI's own caveats (reported): over half of successful 4-8 hour agent tasks needed human intervention, compute also grew, and runtime is not validated research output; no outside audit.
- METR's expenditure horizon (2026-07-21): after more than $10k of agent spend on NanoGPT, agents' horizon was $0-3k versus ~$2.5k of human labor per 1% gain.
- METR judged the evidence in Anthropic's Feb 2026 risk report on R&D automation inadequate, while agreeing the risk was very low for Opus 4.6 (2026-05-08).
- OpenAI says it does not know how to reach fully automated self-improvement safely; the July 2026 Hugging Face breach involved internal agents.
Milestones
- METR pre-deployment summary on Claude Opus 5.5: AI R&D unlikely dramatically accelerated; full automation unlikely METR / Anthropic · 22 September 2026
- OpenAI says it reached its automated-research-intern goal; next target is an automated AI researcher by Mar 2028 OpenAI · 6 September 2026
- METR 'expenditure horizon' measure for AI optimization ability, applied to NanoGPT METR · 21 July 2026
- OpenAI Parameter Golf draws 1,000+ participants and 2,000+ submissions that explore AI-assisted ML research under strict constraints OpenAI · 12 May 2026
- METR reviews the R&D-automation section of Anthropic's Feb 2026 risk report METR / Anthropic · 8 May 2026
- Sakana's AI Scientist published in Nature Sakana AI · 26 March 2026
- Karpathy releases autoresearch, in which agents run LLM-training experiments autonomously Eureka Labs · 6 March 2026
- Altman livestream targets an intern-level research assistant by Sep 2026 and a fully automated AI researcher by 2028 OpenAI · 28 October 2025
- AlphaEvolve, a Gemini-powered evolutionary coding agent, speeds Gemini training and Google data centers Google DeepMind · 14 May 2025
Who is working on it
- Sam Altman, Jakub Pachocki, OpenAI
- Dario Amodei, Anthropic
- Andrej Karpathy (autoresearch), Anthropic (since May 2026; Eureka Labs founder)
- AlphaEvolve team, Google DeepMind
- AI Scientist team, Sakana AI
- Shumeet Baluja, Ian Fischer, Poetiq
The labs with the most milestones here are OpenAI (3), Anthropic (2), METR (1), Eureka Labs (1), Google DeepMind (1) and Sakana AI (1).
Sources
- helpnetsecurity.com/2026/09/07/openai-research-automation-intern/
- runtimewire.com/article/openai-research-intern-devday
- aiunderstanding.org/news/openai-says-coding-agents-now-exceed-human-research-labor-in-its-
- techcrunch.com/2025/10/28/sam-altman-says-openai-will-have-a-legitimate-ai-researcher-by-2
- metr.org/blog/2026-09-22-claude-opus-5-5/
- metr.org/blog/2026-07-21-expenditure-horizon/
- metr.org/blog/2026-05-08-rd-section-anthropic-risk-report-feb-2026-review/
- deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algo
- github.com/karpathy/autoresearch
- sakana.ai/ai-scientist-nature/
- openai.com/news/rss.xml
- openai.com/index/research-acceleration-view-inside-openai/
- sakana.ai/ai-scientist-first-publication/
- techcrunch.com/2026/05/19/openai-co-founder-andrej-karpathy-joins-anthropics-pre-training-
- poetiq.ai
This research bet was checked and corrected against its sources on 6 October 2026. How we check