Automated science and AI scientists

AI agents that propose hypotheses, run experiments in silico or in autonomous labs, and learn from results to make discoveries with little step-by-step direction.

Science and math give verifiable feedback (proofs, assays, simulations) that scales agents beyond text. The variants are evolutionary code search (AlphaEvolve), multi-agent hypothesis systems (co-scientist, Kosmos), Lean-verified math, and lab-in-the-loop models (Periodic, Anthropic's biology lab).

Where it stands. 2026 made math the proving ground, with Claude's Riemann-zeta bound (Aug), a Lean proof of Fermat's Last Theorem (Sep), and Meta and OpenAI open-problem results. Wet-lab discoveries lag, and Anthropic's enzyme has no known function yet.

Evidence

For

  • AlphaEvolve (2025-05-14): on 50+ open math problems it matched best-known results ~75% of the time and improved ~20%, including a 593-sphere kissing configuration; it also sped a Gemini kernel 23%.
  • Claude's Riemann-zeta result (2026-08-10): an unreleased research model raised the proven fraction of zeros on the critical line from 41.6% to 67.2%, checked by Anthropic mathematicians and two outside experts.
  • Claude wrote the first end-to-end Lean proof of Fermat's Last Theorem in 11 days, ~13M lines and 29,500 theorems in the final proof (Anthropic, 2026-09-04); Kevin Buzzard called it a significant step.
  • Meta (2026-10-02): mathematicians used Muse Spark 1.1 and 1.2 in plain chat to write six papers, five answering previously open questions; OpenAI (2026-08-01) also reported results on long-standing open problems.
  • Anthropic (2026-08-18): Claude designed protein binders for 15 targets and succeeded on 14, with 22-35% of designs binding versus 10-15% typical (company-reported).
  • Anthropic (2026-09-23): ~950 Claude agents spent 21 hours and 210M tokens on DNA data to flag a novel CRISPR-like enzyme system; Periodic Labs (2026-09-15) says its lab-trained Neon beats frontier models on X-ray diffraction analysis.

Against

  • Anthropic says Claude did not prove the Riemann hypothesis and its techniques are not expected to; the Fermat result is verification of Wiles's known proof, not new mathematics.
  • Most AI discoveries come where verification is fast (math, for example); experimental fields stay harder and costlier to verify (Anthropic, 2026-08-18).
  • The Anthropic enzyme's function is still unknown; the claim rests on a preprint and in-house lab work.
  • Physical experiments do not scale like digital ones: each takes days, results are ambiguous, and parallel agents are limited by equipment and power (Periodic, 2026-09-15).
  • Vetting is a bottleneck: OpenAI announced an independent Advisory Group on Mathematics and AI (2026-09-21) to advise on review of AI results, and Kosmos's 79.4% statement accuracy means about one in five statements was wrong.

Milestones

Related launches

  • SynthID Bio Google DeepMind · 30 September 2026
    SynthID Bio is a proof of concept for watermarking AI-generated proteins in the sequence itself while preserving function in lab tests.
  • Quine (AI research system for biology) Microsoft · 29 September 2026
    Multimodal biological world model plus orchestration harness that proposes and ranks interventions before wet-lab tests; limited to a fellows programme.
  • Life Sciences Verification Program Anthropic · 17 September 2026
    The Life Sciences Verification Program gives vetted biology teams Mythos, Opus and Sonnet models with biology safeguards loosened, in beta for institutions.
  • Devin-assisted RSA-260 factoring Cognition · 9 September 2026
    Cognition claims Devin helped build a GPU lattice siever that factored the 260-digit RSA-260 challenge in about 4,900 GPU-days (~$400k).
  • AI-generated Navier-Stokes finite-time blow-up proof (with Lean formalization) OpenAI · 8 September 2026
    OpenAI claims an internal agent system proved a Navier-Stokes finite-time singularity with smooth forcing, resolving the Millennium Prize problem, with a Lean proof.
  • AlphaGenome Atlas Google DeepMind · 8 September 2026
    AlphaGenome Atlas gives predicted molecular effects for all 9 billion possible single-letter human DNA variants, a 1 PB resource free for academic use.
  • WeatherNext 3 Google DeepMind · 3 September 2026
    WeatherNext 3 refreshes forecasts hourly from real-time satellite data at up to 5 km resolution and adds precipitation and clean-energy variables.
  • Model Hardware Standard (research preview) Anthropic · 27 August 2026
    The Model Hardware Standard is a shared driver spec letting AI agents operate lab and factory instruments over MCP; research preview opens to select labs.
  • WeatherNext cyclone model (Nature) and open-sourced code and weights Google DeepMind · 6 August 2026
    In a Nature paper, WeatherNext gives cyclone forecasts a day of extra lead (3-day as good as prior 2-day), and Google open-sources the code and weights.
  • Discovering cryptographic weaknesses with Claude Anthropic · 28 July 2026
    Claude Mythos Preview found an improved attack on the post-quantum signature scheme HAWK and a new attack on round-reduced AES.
  • Leanstral 1.5 Mistral AI · 2 July 2026
    Leanstral 1.5 (119B, 6B active) saturates miniF2F at 100% and solves 587 of 672 PutnamBench problems under Apache 2.0.
  • Claude Science Anthropic · 30 June 2026
    Claude Science is a beta workbench where a coordinating agent with 60+ skills and connectors runs reproducible analyses, checked by a reviewer agent.
  • Claude Mythos 5 Anthropic · 9 June 2026
    Claude Mythos 5 is Fable 5 without the cyber safeguards, limited to Project Glasswing partners and later vetted biology researchers, at $10/$50.
  • Disproof of the Erdos planar unit-distance conjecture OpenAI · 20 May 2026
    An unreleased OpenAI reasoning model disproved the 80-year-old belief that square-grid constructions maximize unit distances, with checking by external mathematicians.
  • Gemini for Science Google · 19 May 2026
    Gemini for Science bundles Labs prototypes built on Co-Scientist and AlphaEvolve for hypothesis generation and computational discovery.
  • Intern-S2-Preview Shanghai AI Laboratory (InternLM) · 15 May 2026
    35B scientific model continued from Qwen3.5 that matches the 1T Intern-S1-Pro on core science tasks and generates material crystal structures.
  • AI co-clinician Google DeepMind · 30 April 2026
    AI co-clinician is a DeepMind research initiative for AI agents that work with patients under physician supervision, with zero critical errors in 97 of 98 primary-care evidence…
  • GPT-Rosalind OpenAI · 16 April 2026
    GPT-Rosalind is a frontier reasoning model for biology and drug discovery, available to qualified customers through trusted access.
  • Leanstral Mistral AI · 16 March 2026
    Leanstral, a 120B sparse model with 6B active, is Mistral's open agent for Lean 4 formal proofs, scoring 26.3 at pass@2 for $36.
  • First Proof submissions OpenAI · 20 February 2026
    An internal OpenAI model attempted all 10 First Proof research-level problems; OpenAI judged at least five likely correct and retracted one belief.
  • Single-minus gluon tree amplitudes are nonzero (GPT-5.2) OpenAI · 13 February 2026
    GPT-5.2 proposed a formula for gluon scattering amplitudes that many physicists expected to vanish; an internal model proved it and the authors verified it.
  • Aletheia math research agent (Gemini Deep Think) Google DeepMind · 11 February 2026
    Aletheia, a Deep Think math agent with a natural-language verifier, produced a research paper with no human input and solved four open Erdős-database questions.
  • Intern-S1-Pro Shanghai AI Laboratory (InternLM) · 2 February 2026
    Described as the first one-trillion-parameter scientific multimodal model, with 1T total, 512 experts and 22B active. It is said to master over 100 science tasks.
  • Prism OpenAI · 27 January 2026
    Prism is a free AI-native workspace for scientists to write and collaborate on research, powered by GPT-5.2.
  • Claude for Healthcare Anthropic · 11 January 2026
    Claude for Healthcare launches HIPAA-ready tools with connectors to the CMS Coverage Database, ICD-10 and the NPI registry, plus life-sciences additions.
  • DeepSeekMath-V2 DeepSeek · 27 November 2025
    Open 685B proof model that uses a trained verifier as reward. It reaches gold-level on IMO 2025 and CMO 2024 and 118/120 on Putnam 2024 with scaled test-time compute.
  • WeatherNext 2 Google DeepMind · 17 November 2025
    WeatherNext 2: a Functional Generative Network makes hundreds of coherent forecast scenarios in under a minute on one TPU, beating v1 on 99.9% of variables.
  • AlphaProof methodology in Nature Google DeepMind · 12 November 2025
    Google DeepMind publishes AlphaProof's methodology in Nature. It uses reinforcement learning over formal Lean proofs and is the system behind the 2024 IMO silver.
  • Mathematical exploration and discovery at scale (AlphaEvolve) Google DeepMind · 3 November 2025
    AlphaEvolve is run on 67 problems in analysis, combinatorics, geometry and number theory; it rediscovers most best-known results and improves several.
  • Claude for Life Sciences Anthropic · 20 October 2025
    Claude for Life Sciences adds connectors to Benchling, PubMed, 10x Genomics and others, plus a single-cell RNA QC Agent Skill.
  • Cell2Sentence-Scale 27B (C2S-Scale) Google DeepMind · 15 October 2025
    C2S-Scale is a 27B-parameter Gemma-based single-cell foundation model built with Yale. It flagged a drug combination to make tumors more visible to the immune system.
  • AlphaEarth Foundations Google DeepMind · 30 July 2025
    AlphaEarth Foundations fuses petabytes of satellite and other Earth observation data into per-location embeddings; annual embedding layers released as a dataset.
  • Intern-S1 Shanghai AI Laboratory (InternLM) · 24 July 2025
    241B-total (28B-active) scientific multimodal MoE continually pretrained on 5T tokens (over 2.5T scientific) with Mixture-of-Rewards RL on 1,000+ tasks.
  • Aeneas Google DeepMind · 23 July 2025
    Aeneas helps historians interpret, attribute and restore fragmentary Latin inscriptions by finding parallel texts and proposing restorations; successor to Ithaca.
  • IMO 2025 gold-medal-level result (experimental reasoning model) OpenAI · 19 July 2025
    An experimental OpenAI reasoning model solves 5 of 6 IMO 2025 problems in natural language with no tools, 35/42 points, a gold-medal score.
  • MAI-DxO (Diagnostic Orchestrator) Microsoft · 30 June 2025
    Orchestrator where LLM roles act as a virtual physician panel; with o3 it solves 85.5% of 304 NEJM cases vs 20% for 21 physicians.
  • AlphaGenome Google DeepMind · 25 June 2025
    AlphaGenome reads up to 1 million DNA base pairs and predicts thousands of regulatory properties at single-letter resolution; API for non-commercial research.
  • DeepSeek-Prover-V2 (7B, 671B) DeepSeek · 30 April 2025
    Lean 4 prover built on V3: 88.9% on miniF2F-test and 49 of 658 PutnamBench problems, using V3-driven subgoal decomposition for cold-start data.
  • Qwen2.5-Math-PRM and ProcessBench Alibaba (Qwen) · 14 January 2025
    Qwen releases open process reward models (7B, 72B) and ProcessBench, a 3,400-case benchmark for finding the first erroneous step in math reasoning.
  • GenCast Google DeepMind · 4 December 2024
    GenCast is a diffusion-based ensemble weather model that makes a 15-day probabilistic forecast in 8 minutes on one TPU and beats ECMWF ENS on 97.2% of 1,320 targets.
  • AlphaQubit Google DeepMind · 20 November 2024
    AlphaQubit is a transformer decoder for quantum error correction with 6% fewer errors than tensor networks and 30% fewer than correlated matching.
  • AlphaProteo Google DeepMind · 5 September 2024
    AlphaProteo is an AI designer of de novo protein binders with 3 to 300 times better affinity than prior methods on seven targets, and the first AI binder for VEGF-A.
  • DeepSeek-Prover-V1.5 DeepSeek · 15 August 2024
    Lean 4 prover that adds RL from proof-assistant feedback and an exploration-driven tree search (RMaxTS), reaching 63.5% on miniF2F.
  • Qwen2-Math and Qwen2-Audio Alibaba (Qwen) · 8 August 2024
    Math-specialised Qwen2-Math (2024-08-08) and the voice-chat model Qwen2-Audio (2024-08-09).
  • AlphaProof and AlphaGeometry 2 (IMO 2024 silver) Google DeepMind · 25 July 2024
    AlphaProof plus AlphaGeometry 2 solved 4 of 6 IMO 2024 problems for 28/42 points, silver-medal level.
  • Mathstral 7B Mistral AI · 16 July 2024
    Mathstral, a 7B STEM model built with Project Numina, scores 56.6% on MATH, rising to 74.59% with a reward model.
  • DeepSeek-Prover (V1) DeepSeek · 23 May 2024
    Lean 4 theorem prover trained on 8M synthetic formal statements with proofs, reaching 52% cumulative on miniF2F versus GPT-4's 23%.
  • AlphaFold 3 Google DeepMind · 8 May 2024
    Diffusion-based AlphaFold 3 predicts structures of proteins, DNA, RNA, ligands and ions together, with 50%+ better protein-ligand accuracy than prior methods.
  • Med-Gemini Google · 29 April 2024
    Gemini-based medical models hit 91.1% on MedQA (USMLE) with uncertainty-guided web search; paper 'Capabilities of Gemini Models in Medicine'.
  • DeepSeekMath 7B (introduces GRPO) DeepSeek · 5 February 2024
    7B math model at 51.7% on MATH without tools; introduced GRPO, the critic-free RL algorithm later used for DeepSeek-R1.
  • AlphaGeometry Google DeepMind · 17 January 2024
    Neuro-symbolic geometry solver that solved 25 of 30 olympiad problems, versus 10 for the previous best system.
  • FunSearch Google DeepMind · 14 December 2023
    LLM plus automated evaluator evolves programs; found the largest cap sets in two decades and better bin-packing heuristics.
  • GNoME Google DeepMind · 29 November 2023
    Graph networks predicted 2.2 million new crystals, 380,000 of them stable; 736 later independently synthesized.
  • GraphCast Google DeepMind · 14 November 2023
    Graph-neural-network weather model making a 10-day global forecast in under a minute on one TPU v4, beating ECMWF HRES on most targets.
  • AlphaMissense Google DeepMind · 19 September 2023
    AlphaFold-derived model classifying 89% of all 71 million human missense variants as likely pathogenic or likely benign.
  • Med-PaLM 2 Google · 16 May 2023
    Med-PaLM 2 scored 86.5% on MedQA (USMLE-style), 19+ points above Med-PaLM, with physicians preferring its answers on 8 of 9 axes.

Every launch in the atlas

Who is working on it

The labs with the most milestones and launches here are Google DeepMind (30), Anthropic (11), OpenAI (10), DeepSeek (5), Mistral AI (4) and Shanghai AI Laboratory (InternLM) (3).

Sources

  1. anthropic.com/research/riemann-zeta
  2. anthropic.com/research/formalizing-fermats-last-theorem
  3. anthropic.com/research/Claude-accelerates-protein-design
  4. anthropic.com/news/claude-discovers-novel-enzyme-system
  5. research.meta.ai/blog/solving-open-research-problems-together
  6. openai.com/news/rss.xml
  7. deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algo
  8. research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/
  9. arxiv.org/abs/2511.02824
  10. arxiv.org/abs/2511.16072
  11. nature.com/articles/s41586-025-09833-y
  12. sakana.ai/ai-scientist-nature/
  13. huggingface.co/mistralai/Leanstral-1.5-119B-A6B
  14. periodic.com/news/building-labs-that-learn
  15. openai.com/index/advisory-group-on-mathematics-and-ai/
  16. openai.com/index/ten-advances-in-mathematics/
  17. deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-m

This research bet was checked and corrected against its sources on 6 October 2026. How we check