Reasoning models

Models and papers about reasoning, where a model works through a problem step by step before it answers. The atlas has logged 270 launches and papers on it since 10 May 2023, 106 of them in 2026.

Google DeepMind, OpenAI and Anthropic have the most.

Key launches

The 60 most significant of 270, newest first. Every launch in the atlas

  • Claude Haiku 5.5 Anthropic · 7 October 2026
    Anthropic released Claude Haiku 5.5, its smallest model, with a 1M token context window at $0.10 per million input tokens and $0.50 output, about 75% cheaper than Haiku 4.5 on…
  • Sharing AI progress in mathematics (722 math manuscripts) OpenAI · 6 October 2026
    OpenAI published 722 math manuscripts in 372 result families, produced by an unreleased internal frontier model, in a public GitHub repository with Lean formalizations of many…
  • OpenAI math results release (722 manuscripts, 372 result groups) OpenAI · 6 October 2026
    OpenAI published 722 math manuscripts in 372 groups of related results on GitHub, produced by an unreleased internal model that attempted about 4,000 research problems.
  • 377 solved unsolved math problems OpenAI · 6 October 2026
    OpenAI published more than 700 papers describing solutions to 377 previously unsolved math problems, produced by an unreleased AI model.
  • Beam Reflection AI · 5 October 2026
    Reflection AI unveiled Beam, its first open-weight model, a 501B parameter mixture-of-experts model with 23B active parameters, aimed at coding and agent tasks.
  • Gemini 4 Argon Google DeepMind · 30 September 2026
    Gemini 4 Argon, a new frontier model with a 1M-token output limit, scores 77.9% on DeepSWE v1.1; released first to vetted cyber defenders via Fairwind.
  • GPT-Synopsys Synopsys / OpenAI · 30 September 2026
    Synopsys and OpenAI announced a multi-year partnership to build GPT-Synopsys, a specialized model that operates Synopsys EDA tools to run chip design workflows, with shared…
  • Claude Opus 5.5 Anthropic · 22 September 2026
    Claude Opus 5.5 matches Fable 5.1 on most work, costs 40% less than Opus 5, and scores 66.4% on Terminal-Bench 4.0.
  • GPT-6 Sol and GPT-6 Luna OpenAI · 22 September 2026
    GPT-6 Sol and Luna bring Astra-era training to cheaper tiers at half the price of their GPT-5.6 equivalents.
  • MiMo-V2.6-Pro and V2.6-Flash Xiaomi (MiMo) · 21 September 2026
    Omni-modal 1.02T (Pro) and 310B (Flash) MIT-licensed models; Artificial Analysis rated Pro 46, top of open-weight models at launch.
  • AI-generated Navier-Stokes finite-time blow-up proof (with Lean formalization) OpenAI · 8 September 2026
    OpenAI claims an internal agent system proved a Navier-Stokes finite-time singularity with smooth forcing, resolving the Millennium Prize problem, with a Lean proof.
  • Formalizing Fermat's Last Theorem Anthropic · 4 September 2026
    Claude agents wrote the first complete computer-checked Lean proof of Fermat's Last Theorem in about 11 days, with 13 million lines of Lean.
  • GPT-6 Astra OpenAI · 3 September 2026
    GPT-6 Astra, OpenAI's first model rated Critical for cybersecurity, claims 97.6% on FrontierMath Tier 4 and 99.9% on ARC-AGI-3.
  • Claude Fable 5.1 and Mythos 5.1 Anthropic · 1 September 2026
    Claude Fable 5.1 lifts Terminal-Bench 4.0 from 42.0% to 55.8% and cuts cache-read price 75%; Mythos 5.1 is the same model with fewer safeguards.
  • Hy4 preview Tencent · 28 August 2026
    770B-total, 49B-active open MoE with 1M-token context, Gated DeepSeek Sparse Attention and hyper-connections, under Apache-2.0.
  • Qwen3.8-2.4T-A95B (open weights) Alibaba (Qwen) · 12 August 2026
    Open weights of a 2.4T-parameter, 95B-active MoE: Terminal Bench 2.1 86.6 and SWE-bench Pro 67.7, under a Qwen3.8-Max license.
  • Claude raises the lower bound on Riemann zeta zeros on the critical line Anthropic · 10 August 2026
    An unreleased Claude raised the proven fraction of Riemann zeta zeros on the critical line from 41.6% to 67.2% while attempting the Riemann hypothesis.
  • Qwen3.8-Max Alibaba (Qwen) · 2 August 2026
    Alibaba's largest model has 2.4T parameters (95B active) and 1M context, and is positioned for coding and 'cowork'. Open weights followed the next week.
  • Ten advances in mathematics and theoretical computer science OpenAI · 1 August 2026
    An internal Astra found ten results on decade-old open problems, including non-sofic groups, each with a Lean certificate.
  • Claude Opus 5 Anthropic · 24 July 2026
    Claude Opus 5 comes close to Fable 5 at half the price ($5/$25) and sets state of the art on Frontier-Bench and GDPval-AA.
  • Kimi K3 Moonshot AI · 16 July 2026
    2.8T-parameter open-weights MoE (104B active, 1M context) on Kimi Delta Attention and Attention Residuals, reported third on Artificial Analysis behind two closed models at launch.
  • GPT-5.6 (Sol, Terra, Luna) OpenAI · 9 July 2026
    GPT-5.6 splits into Sol, Terra and Luna tiers, with Sol scoring 53.6 on Agents' Last Exam and 92.2% on BrowseComp at lower token cost.
  • Hy3 Tencent · 6 July 2026
    Completed Hy3: same 295B/21B MoE, now Apache-2.0, API input at 1 yuan per million tokens; scored 2.67 vs GLM-5.1's 2.51 in a 270-expert blind test.
  • Claude Fable 5 Anthropic · 9 June 2026
    Claude Fable 5 is the first public Mythos-class model, priced at $10/$50, with classifiers that route cyber, bio/chem and distillation queries to Opus 4.8.
  • Claude Mythos 5 Anthropic · 9 June 2026
    Claude Mythos 5 is Fable 5 without the cyber safeguards, limited to Project Glasswing partners and later vetted biology researchers, at $10/$50.
  • MiniMax M3 MiniMax · 1 June 2026
    Natively multimodal ~428B-total (~23B active) model with MiniMax Sparse Attention for 1M context at about 1/20 the per-token cost of M2; 80.5% SWE-bench Verified.
  • Disproof of the Erdos planar unit-distance conjecture OpenAI · 20 May 2026
    An unreleased OpenAI reasoning model disproved the 80-year-old belief that square-grid constructions maximize unit distances, with checking by external mathematicians.
  • Gemini 3.5 Flash Google DeepMind · 19 May 2026
    Gemini 3.5 Flash, launched at I/O 2026, beats 3.1 Pro on agentic benchmarks, with 76.2% on Terminal-Bench 2.1, 1656 Elo on GDPval-AA and 83.6% on MCP Atlas.
  • Natural Language Autoencoders Anthropic · 7 May 2026
    Natural Language Autoencoders train Claude to explain its own activations in text, checked by a second copy reconstructing the activation from the explanation.
  • DeepSeek-V4-Pro and V4-Flash (preview) DeepSeek · 24 April 2026
    1.6T-parameter MoE (49B active) plus 284B Flash, 1M-token context; at 1M tokens uses 27% of V3.2's FLOPs and 10% of its KV cache.
  • DeepSeek-V4 (CSA/HCA attention, mHC, Muon) DeepSeek · 24 April 2026
    1.6T-parameter V4-Pro (49B active) and 284B V4-Flash (13B active) with 1M-token context; Pro uses 27% of V3.2's inference FLOPs and 10% of its KV cache.
  • GPT-5.5 and GPT-5.5 Pro OpenAI · 23 April 2026
    GPT-5.5 reaches 82.7% on Terminal-Bench 2.0 and 78.7% on OSWorld-Verified; the API followed a day later with a 1M-token window.
  • gpt-image-2 (ChatGPT Images 2.0) OpenAI · 21 April 2026
    gpt-image-2 / ChatGPT Images 2.0: an image model with built-in thinking (web search, multi-image output, self-checks), up to 2K.
  • Claude Opus 4.7 Anthropic · 16 April 2026
    Claude Opus 4.7 raises SWE-bench Pro to 64.3% and accepts images over 3x larger in pixels (2,576 px), while deliberately limiting cyber capability.
  • Automated Alignment Researchers Anthropic · 14 April 2026
    Claude agents working as automated alignment researchers closed 0.97 of a weak-to-strong supervision gap in five days, versus 0.23 for human researchers in seven.
  • Muse Spark Meta · 8 April 2026
    First Meta Superintelligence Labs model, a closed, natively multimodal reasoning model on a rebuilt stack and Meta's first flagship without open weights.
  • Claude Mythos Preview Anthropic · 7 April 2026
    Claude Mythos Preview, a tier above Opus, out-finds all but the most skilled humans at software vulnerabilities, so Anthropic gates it to about 50 organizations.
  • Gemma 4 (E2B, E4B, 26B MoE, 31B) Google DeepMind · 2 April 2026
    Gemma 4 ships under Apache 2.0 in four sizes; the 31B dense ranked 3 among open models on Arena AI text, the 26B MoE 6.
  • GPT-5.4 and GPT-5.4 Pro OpenAI · 5 March 2026
    GPT-5.4 is OpenAI's first general-purpose model with native computer use, scoring 75.0% on OSWorld-Verified, with a 1M-token context.
  • Gemini 3.1 Pro Google DeepMind · 19 February 2026
    Gemini 3.1 Pro scored a verified 77.1% on ARC-AGI-2, more than double Gemini 3 Pro, with 80.6% on SWE-bench Verified and 94.3% on GPQA Diamond.
  • Qwen3.5-397B-A17B and Qwen3.5-Plus Alibaba (Qwen) · 16 February 2026
    Open 397B-A17B native vision-language MoE with Gated DeltaNet linear attention, 201 languages and 262K context; hosted Qwen3.5-Plus offers 1M.
  • Seed 2.0 (Pro, Lite, Mini, Code) ByteDance Seed · 14 February 2026
    Flagship Doubao generation in Pro, Lite, Mini and Code models; Pro claims IMO, CMO and ICPC gold results and token prices about ten times lower.
  • Gemini 3 Deep Think upgrade (Feb 2026) Google · 12 February 2026
    Upgraded Deep Think, built on Gemini 3.1 Pro, reports 84.6% on ARC-AGI-2 (ARC Prize verified), 48.4% on Humanity's Last Exam and a 3455 Codeforces Elo.
  • GLM-5 Zhipu AI / Z.ai · 11 February 2026
    744B-total (40B active) MIT-licensed MoE with DeepSeek Sparse Attention, trained on 28.5T tokens; 77.8% SWE-bench Verified and top open model on Artificial Analysis at launch.
  • Aletheia math research agent (Gemini Deep Think) Google DeepMind · 11 February 2026
    Aletheia, a Deep Think math agent with a natural-language verifier, produced a research paper with no human input and solved four open Erdős-database questions.
  • Claude Opus 4.6 Anthropic · 5 February 2026
    Claude Opus 4.6 adds a 1M-token context (beta), 128K output, adaptive thinking and Claude Code agent teams at unchanged $5/$25 pricing.
  • Gemini 3 Flash Google DeepMind · 17 December 2025
    Gemini 3 Flash has Pro-grade reasoning at $0.50 / $3 per million tokens. It scores 78% on SWE-bench Verified, above 3 Pro, and is 3x faster than 2.5 Pro.
  • GPT-5.2 OpenAI · 11 December 2025
    GPT-5.2 targets professional knowledge work, reaching 70.9% wins or ties vs experts on GDPval and 52.9% on ARC-AGI-2.
  • Gemini 3 Deep Think Google · 4 December 2025
    Gemini 3 Deep Think reaches Google AI Ultra subscribers with 41.0% on Humanity's Last Exam and 45.1% on ARC-AGI-2 (with code execution, ARC Prize verified).
  • DeepSeek-V3.2 (DSA, scaled RL, agentic synthesis) DeepSeek · 2 December 2025
    Pairs sparse attention with a scaled RL budget and a synthetic agentic-task pipeline; V3.2-Speciale claims IMO and IOI 2025 gold-level results.
  • Mistral 3: Mistral Large 3 and Ministral 3 Mistral AI · 2 December 2025
    Mistral Large 3, a 675B-parameter MoE with 41B active trained on 3,000 H200s, and Ministral 3 (3B/8B/14B) all ship under Apache 2.0.
  • Gemini 3 Pro Google DeepMind · 18 November 2025
    Gemini 3 Pro scored 1501 Elo on LMArena, 37.5% on Humanity's Last Exam, 91.9% on GPQA Diamond and 76.2% on SWE-bench Verified, and shipped in Search on day one.
  • Gemini Deep Think at IMO 2025 (gold-medal score) Google DeepMind · 21 July 2025
    Advanced Gemini Deep Think scored 35/42 at IMO 2025, a gold-medal score, working end to end in natural language within 4.5 hours; IMO-graded.
  • DeepSeek-R1: Incentivizing Reasoning via RL (arXiv) DeepSeek · 22 January 2025
    Shows reasoning can emerge from pure RL on verifiable rewards (R1-Zero), then builds R1 with a small cold-start and distills it to small models.
  • DeepSeek-R1 and R1-Zero DeepSeek · 20 January 2025
    Open-weights 671B MoE reasoning model claimed at OpenAI o1 level; R1-Zero showed reasoning emerging from pure RL with rule-based rewards.
  • o3 (announcement) OpenAI · 20 December 2024
    o3 scores 75.7% (high-efficiency) and 87.5% (low-efficiency) on ARC-AGI-1 semi-private tasks; announced on day 12 of OpenAI's December event, not released.
  • o1-preview and o1-mini OpenAI · 12 September 2024
    First model trained with large-scale reinforcement learning to think in a long hidden chain of thought before answering.
  • Gemini 1.5 Pro Google DeepMind · 15 February 2024
    Mixture-of-experts Gemini 1.5 Pro matched 1.0 Ultra with less compute and offered a 1M-token context (10M tested in research).
  • DeepSeekMath (introduces GRPO) DeepSeek · 5 February 2024
    Introduces Group Relative Policy Optimization (GRPO), a critic-free PPO variant, plus a 7B math model scoring 51.7% on MATH.
  • Gemini 1.0 (Ultra, Pro, Nano) Google DeepMind · 6 December 2023
    First Gemini family, natively multimodal, with Ultra, Pro and Nano models; Ultra reported 90.0% on MMLU with CoT@32, beating GPT-4's reported score.

Most active labs