Automated AI R&D (recursive self-improvement)

Use AI agents to do AI research itself; OpenAI says it has an 'automated research intern' and targets an automated AI researcher by March 2028.

If agents can design experiments, write code and choose ideas, research speed compounds and capability growth accelerates. OpenAI reports 3.1 agent-workdays per human workday inside its research org; others pursue it through evolutionary code search, AI-scientist loops and autoresearch.

Where it stands. OpenAI claims 3.1 agent-workdays per human workday; METR judges AI R&D at Anthropic unlikely to be dramatically accelerated so far. OpenAI's next declared milestone is an automated AI researcher by March 2028.

Evidence

For

  • OpenAI (2026-09-06): its research org logged 3.1 agent-workdays per human workday by mid-August; the median researcher spends $600+/day on inference at API prices (company-reported).
  • AlphaEvolve (2025-05-14) sped a Gemini kernel 23%, cutting Gemini training time 1%, and a Borg heuristic recovers ~0.7% of Google's fleet compute.
  • Self-reported researcher uplift: ~4x in Anthropic's Mythos Preview system card and ~2x in METR's Becker et al. (2026), both flagged as unreliable (METR, 2026-07-21).
  • Karpathy's autoresearch (2026-03-06): an agent edits training code and runs 5-minute experiments overnight, an open template for agent-run research orgs.
  • Sakana's AI Scientist is published in Nature (2026-03-26) after an AI-written paper passed ICLR-workshop review (2025-03-12).
  • A preliminary METR report puts Anthropic's AI-driven acceleration near 1.5x, with a 30% chance of 2x (cited 2026-09-22).

Against

  • METR's Opus 5.5 evaluation (2026-09-22) judged AI R&D at Anthropic unlikely to have been dramatically accelerated and full automation of AI R&D unlikely for this model; part of its evidence came from a separate team that could not share details, plus a source it could not disclose.
  • OpenAI's own caveats (reported): over half of successful 4-8 hour agent tasks needed human intervention, compute also grew, and runtime is not validated research output; no outside audit.
  • METR's expenditure horizon (2026-07-21): after more than $10k of agent spend on NanoGPT, agents' horizon was $0-3k versus ~$2.5k of human labor per 1% gain.
  • METR judged the evidence in Anthropic's Feb 2026 risk report on R&D automation inadequate, while agreeing the risk was very low for Opus 4.6 (2026-05-08).
  • OpenAI says it does not know how to reach fully automated self-improvement safely; the July 2026 Hugging Face breach involved internal agents.

Milestones

Who is working on it

  • Sam Altman, Jakub Pachocki, OpenAI
  • Dario Amodei, Anthropic
  • Andrej Karpathy (autoresearch), Anthropic (since May 2026; Eureka Labs founder)
  • AlphaEvolve team, Google DeepMind
  • AI Scientist team, Sakana AI
  • Shumeet Baluja, Ian Fischer, Poetiq

The labs with the most milestones here are OpenAI (3), Anthropic (2), METR (1), Eureka Labs (1), Google DeepMind (1) and Sakana AI (1).

Sources

  1. helpnetsecurity.com/2026/09/07/openai-research-automation-intern/
  2. runtimewire.com/article/openai-research-intern-devday
  3. aiunderstanding.org/news/openai-says-coding-agents-now-exceed-human-research-labor-in-its-
  4. techcrunch.com/2025/10/28/sam-altman-says-openai-will-have-a-legitimate-ai-researcher-by-2
  5. metr.org/blog/2026-09-22-claude-opus-5-5/
  6. metr.org/blog/2026-07-21-expenditure-horizon/
  7. metr.org/blog/2026-05-08-rd-section-anthropic-risk-report-feb-2026-review/
  8. deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algo
  9. github.com/karpathy/autoresearch
  10. sakana.ai/ai-scientist-nature/
  11. openai.com/news/rss.xml
  12. openai.com/index/research-acceleration-view-inside-openai/
  13. sakana.ai/ai-scientist-first-publication/
  14. techcrunch.com/2026/05/19/openai-co-founder-andrej-karpathy-joins-anthropics-pre-training-
  15. poetiq.ai

This research bet was checked and corrected against its sources on 6 October 2026. How we check