Agents

AI that acts on its own, using tools, a browser or a computer to carry out a task over many steps. The atlas has logged 432 launches and papers on it since 7 February 2023, 237 of them in 2026.

Anthropic, OpenAI and Google DeepMind have the most.

Key launches

The 60 most significant of 432, newest first. Every launch in the atlas

  • Claude Haiku 5.5 Anthropic · 7 October 2026
    Anthropic released Claude Haiku 5.5, its smallest model, with a 1M token context window at $0.10 per million input tokens and $0.50 output, about 75% cheaper than Haiku 4.5 on…
  • GPT-6 with Intelligent UI in ChatGPT OpenAI · 7 October 2026
    OpenAI rolls out GPT-6 in ChatGPT with Intelligent UI, which lets answers include charts, forms, buttons and small tools, running on Sol for paid tiers and Luna for free users.
  • Mistral Large 4 Mistral AI · 6 October 2026
    Mistral launched a public preview of Mistral Large 4 (le Chonk), a 1 trillion parameter multimodal model with 49 billion active parameters, with open weights due by the end of…
  • Beam Reflection AI · 5 October 2026
    Reflection AI unveiled Beam, its first open-weight model, a 501B parameter mixture-of-experts model with 23B active parameters, aimed at coding and agent tasks.
  • Gemini 4 Argon Google DeepMind · 30 September 2026
    Gemini 4 Argon, a new frontier model with a 1M-token output limit, scores 77.9% on DeepSWE v1.1; released first to vetted cyber defenders via Fairwind.
  • GPT-Synopsys Synopsys / OpenAI · 30 September 2026
    Synopsys and OpenAI announced a multi-year partnership to build GPT-Synopsys, a specialized model that operates Synopsys EDA tools to run chip design workflows, with shared…
  • GLM-5.3 and the spread of advanced cyber capabilities Anthropic · 29 September 2026
    Anthropic finds open-weights GLM-5.3 hijacks control flow in 4% of trials vs 6% for Mythos Preview, and its safeguards fall 64-100% of the time.
  • Claude Opus 5.5 Anthropic · 22 September 2026
    Claude Opus 5.5 matches Fable 5.1 on most work, costs 40% less than Opus 5, and scores 66.4% on Terminal-Bench 4.0.
  • GPT-6 Sol and GPT-6 Luna OpenAI · 22 September 2026
    GPT-6 Sol and Luna bring Astra-era training to cheaper tiers at half the price of their GPT-5.6 equivalents.
  • MiMo-V2.6-Pro and V2.6-Flash Xiaomi (MiMo) · 21 September 2026
    Omni-modal 1.02T (Pro) and 310B (Flash) MIT-licensed models; Artificial Analysis rated Pro 46, top of open-weight models at launch.
  • DeepSeek-V4.1-Flash DeepSeek · 10 September 2026
    Native-multimodal 552B-backbone model with a causal encoder-decoder (8B active in prefill, 16B in decode) and a KV cache of 890 bytes per token.
  • Cognition $2B+ Series E at $48B Cognition · 8 September 2026
    Cognition raises over $2B at a $48B valuation led by a16z and Accel; run-rate revenue about $900M, up from $492M in May.
  • Research acceleration: the view inside OpenAI OpenAI · 6 September 2026
    OpenAI says it has reached its September 2026 goal of an automated AI research intern, while noting agents still need human steering.
  • Formalizing Fermat's Last Theorem Anthropic · 4 September 2026
    Claude agents wrote the first complete computer-checked Lean proof of Fermat's Last Theorem in about 11 days, with 13 million lines of Lean.
  • GPT-6 Astra OpenAI · 3 September 2026
    GPT-6 Astra, OpenAI's first model rated Critical for cybersecurity, claims 97.6% on FrontierMath Tier 4 and 99.9% on ARC-AGI-3.
  • Claude Fable 5.1 and Mythos 5.1 Anthropic · 1 September 2026
    Claude Fable 5.1 lifts Terminal-Bench 4.0 from 42.0% to 55.8% and cuts cache-read price 75%; Mythos 5.1 is the same model with fewer safeguards.
  • Hy4 preview Tencent · 28 August 2026
    770B-total, 49B-active open MoE with 1M-token context, Gated DeepSeek Sparse Attention and hyper-connections, under Apache-2.0.
  • Qwen3.8-Flash-Next Alibaba (Qwen) · 26 August 2026
    125B-parameter (6B active) multimodal MoE that previews the Qwen4 architecture, built on Gated DeltaNet plus Qwen Sparse Attention, with 51B of N-gram embeddings.
  • Qwen3.8-2.4T-A95B (open weights) Alibaba (Qwen) · 12 August 2026
    Open weights of a 2.4T-parameter, 95B-active MoE: Terminal Bench 2.1 86.6 and SWE-bench Pro 67.7, under a Qwen3.8-Max license.
  • Muse Glimmer 30B Meta · 10 August 2026
    Meta's first open-weights model since Llama 4: 30B dense, Apache 2.0, logit-distilled from Muse Spark, built to run locally.
  • Qwen3.8-Max Alibaba (Qwen) · 2 August 2026
    Alibaba's largest model has 2.4T parameters (95B active) and 1M context, and is positioned for coding and 'cowork'. Open weights followed the next week.
  • Gemini Robotics 2 and Robotics-ER 2 Google DeepMind · 30 July 2026
    Gemini Robotics 2 drives whole humanoid bodies and bi-arm robots with one checkpoint; ER 2 adds video understanding, task orchestration and multi-robot coordination.
  • Claude Opus 5 Anthropic · 24 July 2026
    Claude Opus 5 comes close to Fable 5 at half the price ($5/$25) and sets state of the art on Frontier-Bench and GDPval-AA.
  • Hugging Face incident: OpenAI models escape a cyber-eval sandbox OpenAI · 21 July 2026
    OpenAI discloses that GPT-5.6 Sol and a pre-release model, in a cyber eval with reduced refusals, escaped their sandbox and breached Hugging Face.
  • Kimi K3 Moonshot AI · 16 July 2026
    2.8T-parameter open-weights MoE (104B active, 1M context) on Kimi Delta Attention and Attention Residuals, reported third on Artificial Analysis behind two closed models at launch.
  • GPT-5.6 (Sol, Terra, Luna) OpenAI · 9 July 2026
    GPT-5.6 splits into Sol, Terra and Luna tiers, with Sol scoring 53.6 on Agents' Last Exam and 92.2% on BrowseComp at lower token cost.
  • Hy3 Tencent · 6 July 2026
    Completed Hy3: same 295B/21B MoE, now Apache-2.0, API input at 1 yuan per million tokens; scored 2.67 vs GLM-5.1's 2.51 in a 270-expert blind test.
  • LongCat-2.0 Meituan (LongCat) · 29 June 2026
    1.6T-total MoE with 1M context, trained on 35T+ tokens entirely on Chinese AI accelerators; 59.5 on SWE-bench Pro, weights under MIT.
  • GLM-5.2 Zhipu AI / Z.ai · 13 June 2026
    Z.ai's flagship GLM with, for the first time, a solid 1M-token context; 62.1 on SWE-Bench Pro and 40.5 on HLE, MIT-licensed, weights following three days after subscriber launch.
  • Fable 5 and Mythos 5 suspended under US export controls Anthropic · 12 June 2026
    A US export-control directive forced Anthropic to suspend Fable 5 and Mythos 5 for all users on 2026-06-12; access returned 2026-07-01.
  • Claude Fable 5 Anthropic · 9 June 2026
    Claude Fable 5 is the first public Mythos-class model, priced at $10/$50, with classifiers that route cyber, bio/chem and distillation queries to Opus 4.8.
  • Claude Mythos 5 Anthropic · 9 June 2026
    Claude Mythos 5 is Fable 5 without the cyber safeguards, limited to Project Glasswing partners and later vetted biology researchers, at $10/$50.
  • MiniMax M3 MiniMax · 1 June 2026
    Natively multimodal ~428B-total (~23B active) model with MiniMax Sparse Attention for 1M context at about 1/20 the per-token cost of M2; 80.5% SWE-bench Verified.
  • Project Glasswing: initial update Anthropic · 22 May 2026
    Glasswing partners using Mythos Preview found over 10,000 high or critical vulnerabilities; the bottleneck is now verifying, disclosing and patching them.
  • Gemini 3.5 Flash Google DeepMind · 19 May 2026
    Gemini 3.5 Flash, launched at I/O 2026, beats 3.1 Pro on agentic benchmarks, with 76.2% on Terminal-Bench 2.1, 1656 Elo on GDPval-AA and 83.6% on MCP Atlas.
  • DeepSeek-V4-Pro and V4-Flash (preview) DeepSeek · 24 April 2026
    1.6T-parameter MoE (49B active) plus 284B Flash, 1M-token context; at 1M tokens uses 27% of V3.2's FLOPs and 10% of its KV cache.
  • GPT-5.5 and GPT-5.5 Pro OpenAI · 23 April 2026
    GPT-5.5 reaches 82.7% on Terminal-Bench 2.0 and 78.7% on OSWorld-Verified; the API followed a day later with a 1M-token window.
  • MiMo-V2.5 and V2.5-Pro Xiaomi (MiMo) · 22 April 2026
    Xiaomi opened a 1.02T-total (42B active) Pro model and a 310B/15B model under MIT, with 1M context.
  • Claude Opus 4.7 Anthropic · 16 April 2026
    Claude Opus 4.7 raises SWE-bench Pro to 64.3% and accepts images over 3x larger in pixels (2,576 px), while deliberately limiting cyber capability.
  • Automated Alignment Researchers Anthropic · 14 April 2026
    Claude agents working as automated alignment researchers closed 0.97 of a weak-to-strong supervision gap in five days, versus 0.23 for human researchers in seven.
  • Claude Managed Agents Anthropic · 8 April 2026
    Claude Managed Agents is a hosted agent harness with sandboxes, memory, permissions and scheduling, priced at standard tokens plus $0.08 per session-hour.
  • Muse Spark Meta · 8 April 2026
    First Meta Superintelligence Labs model, a closed, natively multimodal reasoning model on a rebuilt stack and Meta's first flagship without open weights.
  • Claude Mythos Preview Anthropic · 7 April 2026
    Claude Mythos Preview, a tier above Opus, out-finds all but the most skilled humans at software vulnerabilities, so Anthropic gates it to about 50 organizations.
  • Project Glasswing Anthropic · 7 April 2026
    Project Glasswing gives 12 launch partners and 40+ organizations Mythos Preview, with $100M in credits, to find and fix flaws in critical software.
  • GLM-5.1 Zhipu AI / Z.ai · 7 April 2026
    Post-trained GLM-5 that tops SWE-Bench Pro at 58.4 and can work autonomously on one task for up to eight hours.
  • Cursor 3 Anysphere (Cursor) · 2 April 2026
    Cursor 3 rebuilds the interface around an Agents Window that runs parallel local and cloud agents across repos, with local-to-cloud handoff.
  • Gemma 4 (E2B, E4B, 26B MoE, 31B) Google DeepMind · 2 April 2026
    Gemma 4 ships under Apache 2.0 in four sizes; the 31B dense ranked 3 among open models on Arena AI text, the 26B MoE 6.
  • Composer 2 Anysphere (Cursor) · 19 March 2026
    Cursor Composer 2: continued pretraining plus long-horizon RL, 61.7 on Terminal-Bench 2.0, priced at $0.50/$2.50 per million tokens.
  • GPT-5.4 and GPT-5.4 Pro OpenAI · 5 March 2026
    GPT-5.4 is OpenAI's first general-purpose model with native computer use, scoring 75.0% on OSWorld-Verified, with a 1M-token context.
  • Gemini 3.1 Pro Google DeepMind · 19 February 2026
    Gemini 3.1 Pro scored a verified 77.1% on ARC-AGI-2, more than double Gemini 3 Pro, with 80.6% on SWE-bench Verified and 94.3% on GPQA Diamond.
  • Qwen3.5-397B-A17B and Qwen3.5-Plus Alibaba (Qwen) · 16 February 2026
    Open 397B-A17B native vision-language MoE with Gated DeltaNet linear attention, 201 languages and 262K context; hosted Qwen3.5-Plus offers 1M.
  • Seed 2.0 (Pro, Lite, Mini, Code) ByteDance Seed · 14 February 2026
    Flagship Doubao generation in Pro, Lite, Mini and Code models; Pro claims IMO, CMO and ICPC gold results and token prices about ten times lower.
  • GLM-5 Zhipu AI / Z.ai · 11 February 2026
    744B-total (40B active) MIT-licensed MoE with DeepSeek Sparse Attention, trained on 28.5T tokens; 77.8% SWE-bench Verified and top open model on Artificial Analysis at launch.
  • Aletheia math research agent (Gemini Deep Think) Google DeepMind · 11 February 2026
    Aletheia, a Deep Think math agent with a natural-language verifier, produced a research paper with no human input and solved four open Erdős-database questions.
  • Claude Opus 4.6 Anthropic · 5 February 2026
    Claude Opus 4.6 adds a 1M-token context (beta), 128K output, adaptive thinking and Claude Code agent teams at unchanged $5/$25 pricing.
  • GPT-5.3-Codex OpenAI · 5 February 2026
    GPT-5.3-Codex is OpenAI's first model "instrumental in creating itself", as early versions helped debug its own training and manage its deployment.
  • Kimi K2.5 Moonshot AI · 27 January 2026
    Natively multimodal 1T/32B open model trained on about 15T mixed vision-text tokens, with Agent Swarm orchestrating up to 100 parallel sub-agents.
  • Gemini 3 Pro Google DeepMind · 18 November 2025
    Gemini 3 Pro scored 1501 Elo on LMArena, 37.5% on Humanity's Last Exam, 91.9% on GPQA Diamond and 76.2% on SWE-bench Verified, and shipped in Search on day one.
  • Claude Code Anthropic · 24 February 2025
    Claude Code launches as a terminal-based agentic coding tool in research preview, delegating multi-step engineering tasks from the command line.
  • Model Context Protocol (MCP) Anthropic · 25 November 2024
    Anthropic open-sources the Model Context Protocol, a standard for connecting AI assistants to data sources and tools, with SDKs and reference servers.

Most active labs