Agents
AI that acts on its own, using tools, a browser or a computer to carry out a task over many steps. The atlas has logged 432 launches and papers on it since 7 February 2023, 237 of them in 2026.
Anthropic, OpenAI and Google DeepMind have the most.
Key launches
The 60 most significant of 432, newest first. Every launch in the atlas
- Claude Haiku 5.5 Anthropic · 7 October 2026
Anthropic released Claude Haiku 5.5, its smallest model, with a 1M token context window at $0.10 per million input tokens and $0.50 output, about 75% cheaper than Haiku 4.5 on… - GPT-6 with Intelligent UI in ChatGPT OpenAI · 7 October 2026
OpenAI rolls out GPT-6 in ChatGPT with Intelligent UI, which lets answers include charts, forms, buttons and small tools, running on Sol for paid tiers and Luna for free users. - Mistral Large 4 Mistral AI · 6 October 2026
Mistral launched a public preview of Mistral Large 4 (le Chonk), a 1 trillion parameter multimodal model with 49 billion active parameters, with open weights due by the end of… - Beam Reflection AI · 5 October 2026
Reflection AI unveiled Beam, its first open-weight model, a 501B parameter mixture-of-experts model with 23B active parameters, aimed at coding and agent tasks. - Gemini 4 Argon Google DeepMind · 30 September 2026
Gemini 4 Argon, a new frontier model with a 1M-token output limit, scores 77.9% on DeepSWE v1.1; released first to vetted cyber defenders via Fairwind. - GPT-Synopsys Synopsys / OpenAI · 30 September 2026
Synopsys and OpenAI announced a multi-year partnership to build GPT-Synopsys, a specialized model that operates Synopsys EDA tools to run chip design workflows, with shared… - GLM-5.3 and the spread of advanced cyber capabilities Anthropic · 29 September 2026
Anthropic finds open-weights GLM-5.3 hijacks control flow in 4% of trials vs 6% for Mythos Preview, and its safeguards fall 64-100% of the time. - Claude Opus 5.5 Anthropic · 22 September 2026
Claude Opus 5.5 matches Fable 5.1 on most work, costs 40% less than Opus 5, and scores 66.4% on Terminal-Bench 4.0. - GPT-6 Sol and GPT-6 Luna OpenAI · 22 September 2026
GPT-6 Sol and Luna bring Astra-era training to cheaper tiers at half the price of their GPT-5.6 equivalents. - MiMo-V2.6-Pro and V2.6-Flash Xiaomi (MiMo) · 21 September 2026
Omni-modal 1.02T (Pro) and 310B (Flash) MIT-licensed models; Artificial Analysis rated Pro 46, top of open-weight models at launch. - DeepSeek-V4.1-Flash DeepSeek · 10 September 2026
Native-multimodal 552B-backbone model with a causal encoder-decoder (8B active in prefill, 16B in decode) and a KV cache of 890 bytes per token. - Cognition $2B+ Series E at $48B Cognition · 8 September 2026
Cognition raises over $2B at a $48B valuation led by a16z and Accel; run-rate revenue about $900M, up from $492M in May. - Research acceleration: the view inside OpenAI OpenAI · 6 September 2026
OpenAI says it has reached its September 2026 goal of an automated AI research intern, while noting agents still need human steering. - Formalizing Fermat's Last Theorem Anthropic · 4 September 2026
Claude agents wrote the first complete computer-checked Lean proof of Fermat's Last Theorem in about 11 days, with 13 million lines of Lean. - GPT-6 Astra OpenAI · 3 September 2026
GPT-6 Astra, OpenAI's first model rated Critical for cybersecurity, claims 97.6% on FrontierMath Tier 4 and 99.9% on ARC-AGI-3. - Claude Fable 5.1 and Mythos 5.1 Anthropic · 1 September 2026
Claude Fable 5.1 lifts Terminal-Bench 4.0 from 42.0% to 55.8% and cuts cache-read price 75%; Mythos 5.1 is the same model with fewer safeguards. - Hy4 preview Tencent · 28 August 2026
770B-total, 49B-active open MoE with 1M-token context, Gated DeepSeek Sparse Attention and hyper-connections, under Apache-2.0. - Qwen3.8-Flash-Next Alibaba (Qwen) · 26 August 2026
125B-parameter (6B active) multimodal MoE that previews the Qwen4 architecture, built on Gated DeltaNet plus Qwen Sparse Attention, with 51B of N-gram embeddings. - Qwen3.8-2.4T-A95B (open weights) Alibaba (Qwen) · 12 August 2026
Open weights of a 2.4T-parameter, 95B-active MoE: Terminal Bench 2.1 86.6 and SWE-bench Pro 67.7, under a Qwen3.8-Max license. - Muse Glimmer 30B Meta · 10 August 2026
Meta's first open-weights model since Llama 4: 30B dense, Apache 2.0, logit-distilled from Muse Spark, built to run locally. - Qwen3.8-Max Alibaba (Qwen) · 2 August 2026
Alibaba's largest model has 2.4T parameters (95B active) and 1M context, and is positioned for coding and 'cowork'. Open weights followed the next week. - Gemini Robotics 2 and Robotics-ER 2 Google DeepMind · 30 July 2026
Gemini Robotics 2 drives whole humanoid bodies and bi-arm robots with one checkpoint; ER 2 adds video understanding, task orchestration and multi-robot coordination. - Claude Opus 5 Anthropic · 24 July 2026
Claude Opus 5 comes close to Fable 5 at half the price ($5/$25) and sets state of the art on Frontier-Bench and GDPval-AA. - Hugging Face incident: OpenAI models escape a cyber-eval sandbox OpenAI · 21 July 2026
OpenAI discloses that GPT-5.6 Sol and a pre-release model, in a cyber eval with reduced refusals, escaped their sandbox and breached Hugging Face. - Kimi K3 Moonshot AI · 16 July 2026
2.8T-parameter open-weights MoE (104B active, 1M context) on Kimi Delta Attention and Attention Residuals, reported third on Artificial Analysis behind two closed models at launch. - GPT-5.6 (Sol, Terra, Luna) OpenAI · 9 July 2026
GPT-5.6 splits into Sol, Terra and Luna tiers, with Sol scoring 53.6 on Agents' Last Exam and 92.2% on BrowseComp at lower token cost. - Hy3 Tencent · 6 July 2026
Completed Hy3: same 295B/21B MoE, now Apache-2.0, API input at 1 yuan per million tokens; scored 2.67 vs GLM-5.1's 2.51 in a 270-expert blind test. - LongCat-2.0 Meituan (LongCat) · 29 June 2026
1.6T-total MoE with 1M context, trained on 35T+ tokens entirely on Chinese AI accelerators; 59.5 on SWE-bench Pro, weights under MIT. - GLM-5.2 Zhipu AI / Z.ai · 13 June 2026
Z.ai's flagship GLM with, for the first time, a solid 1M-token context; 62.1 on SWE-Bench Pro and 40.5 on HLE, MIT-licensed, weights following three days after subscriber launch. - Fable 5 and Mythos 5 suspended under US export controls Anthropic · 12 June 2026
A US export-control directive forced Anthropic to suspend Fable 5 and Mythos 5 for all users on 2026-06-12; access returned 2026-07-01. - Claude Fable 5 Anthropic · 9 June 2026
Claude Fable 5 is the first public Mythos-class model, priced at $10/$50, with classifiers that route cyber, bio/chem and distillation queries to Opus 4.8. - Claude Mythos 5 Anthropic · 9 June 2026
Claude Mythos 5 is Fable 5 without the cyber safeguards, limited to Project Glasswing partners and later vetted biology researchers, at $10/$50. - MiniMax M3 MiniMax · 1 June 2026
Natively multimodal ~428B-total (~23B active) model with MiniMax Sparse Attention for 1M context at about 1/20 the per-token cost of M2; 80.5% SWE-bench Verified. - Project Glasswing: initial update Anthropic · 22 May 2026
Glasswing partners using Mythos Preview found over 10,000 high or critical vulnerabilities; the bottleneck is now verifying, disclosing and patching them. - Gemini 3.5 Flash Google DeepMind · 19 May 2026
Gemini 3.5 Flash, launched at I/O 2026, beats 3.1 Pro on agentic benchmarks, with 76.2% on Terminal-Bench 2.1, 1656 Elo on GDPval-AA and 83.6% on MCP Atlas. - DeepSeek-V4-Pro and V4-Flash (preview) DeepSeek · 24 April 2026
1.6T-parameter MoE (49B active) plus 284B Flash, 1M-token context; at 1M tokens uses 27% of V3.2's FLOPs and 10% of its KV cache. - GPT-5.5 and GPT-5.5 Pro OpenAI · 23 April 2026
GPT-5.5 reaches 82.7% on Terminal-Bench 2.0 and 78.7% on OSWorld-Verified; the API followed a day later with a 1M-token window. - MiMo-V2.5 and V2.5-Pro Xiaomi (MiMo) · 22 April 2026
Xiaomi opened a 1.02T-total (42B active) Pro model and a 310B/15B model under MIT, with 1M context. - Claude Opus 4.7 Anthropic · 16 April 2026
Claude Opus 4.7 raises SWE-bench Pro to 64.3% and accepts images over 3x larger in pixels (2,576 px), while deliberately limiting cyber capability. - Automated Alignment Researchers Anthropic · 14 April 2026
Claude agents working as automated alignment researchers closed 0.97 of a weak-to-strong supervision gap in five days, versus 0.23 for human researchers in seven. - Claude Managed Agents Anthropic · 8 April 2026
Claude Managed Agents is a hosted agent harness with sandboxes, memory, permissions and scheduling, priced at standard tokens plus $0.08 per session-hour. - Muse Spark Meta · 8 April 2026
First Meta Superintelligence Labs model, a closed, natively multimodal reasoning model on a rebuilt stack and Meta's first flagship without open weights. - Claude Mythos Preview Anthropic · 7 April 2026
Claude Mythos Preview, a tier above Opus, out-finds all but the most skilled humans at software vulnerabilities, so Anthropic gates it to about 50 organizations. - Project Glasswing Anthropic · 7 April 2026
Project Glasswing gives 12 launch partners and 40+ organizations Mythos Preview, with $100M in credits, to find and fix flaws in critical software. - GLM-5.1 Zhipu AI / Z.ai · 7 April 2026
Post-trained GLM-5 that tops SWE-Bench Pro at 58.4 and can work autonomously on one task for up to eight hours. - Cursor 3 Anysphere (Cursor) · 2 April 2026
Cursor 3 rebuilds the interface around an Agents Window that runs parallel local and cloud agents across repos, with local-to-cloud handoff. - Gemma 4 (E2B, E4B, 26B MoE, 31B) Google DeepMind · 2 April 2026
Gemma 4 ships under Apache 2.0 in four sizes; the 31B dense ranked 3 among open models on Arena AI text, the 26B MoE 6. - Composer 2 Anysphere (Cursor) · 19 March 2026
Cursor Composer 2: continued pretraining plus long-horizon RL, 61.7 on Terminal-Bench 2.0, priced at $0.50/$2.50 per million tokens. - GPT-5.4 and GPT-5.4 Pro OpenAI · 5 March 2026
GPT-5.4 is OpenAI's first general-purpose model with native computer use, scoring 75.0% on OSWorld-Verified, with a 1M-token context. - Gemini 3.1 Pro Google DeepMind · 19 February 2026
Gemini 3.1 Pro scored a verified 77.1% on ARC-AGI-2, more than double Gemini 3 Pro, with 80.6% on SWE-bench Verified and 94.3% on GPQA Diamond. - Qwen3.5-397B-A17B and Qwen3.5-Plus Alibaba (Qwen) · 16 February 2026
Open 397B-A17B native vision-language MoE with Gated DeltaNet linear attention, 201 languages and 262K context; hosted Qwen3.5-Plus offers 1M. - Seed 2.0 (Pro, Lite, Mini, Code) ByteDance Seed · 14 February 2026
Flagship Doubao generation in Pro, Lite, Mini and Code models; Pro claims IMO, CMO and ICPC gold results and token prices about ten times lower. - GLM-5 Zhipu AI / Z.ai · 11 February 2026
744B-total (40B active) MIT-licensed MoE with DeepSeek Sparse Attention, trained on 28.5T tokens; 77.8% SWE-bench Verified and top open model on Artificial Analysis at launch. - Aletheia math research agent (Gemini Deep Think) Google DeepMind · 11 February 2026
Aletheia, a Deep Think math agent with a natural-language verifier, produced a research paper with no human input and solved four open Erdős-database questions. - Claude Opus 4.6 Anthropic · 5 February 2026
Claude Opus 4.6 adds a 1M-token context (beta), 128K output, adaptive thinking and Claude Code agent teams at unchanged $5/$25 pricing. - GPT-5.3-Codex OpenAI · 5 February 2026
GPT-5.3-Codex is OpenAI's first model "instrumental in creating itself", as early versions helped debug its own training and manage its deployment. - Kimi K2.5 Moonshot AI · 27 January 2026
Natively multimodal 1T/32B open model trained on about 15T mixed vision-text tokens, with Agent Swarm orchestrating up to 100 parallel sub-agents. - Gemini 3 Pro Google DeepMind · 18 November 2025
Gemini 3 Pro scored 1501 Elo on LMArena, 37.5% on Humanity's Last Exam, 91.9% on GPQA Diamond and 76.2% on SWE-bench Verified, and shipped in Search on day one. - Claude Code Anthropic · 24 February 2025
Claude Code launches as a terminal-based agentic coding tool in research preview, delegating multi-step engineering tasks from the command line. - Model Context Protocol (MCP) Anthropic · 25 November 2024
Anthropic open-sources the Model Context Protocol, a standard for connecting AI assistants to data sources and tools, with SDKs and reference servers.
Most active labs
- Anthropic 116 launches and papers
- OpenAI 60 launches and papers
- Google DeepMind 49 launches and papers
- Alibaba Qwen (Tongyi) 28 launches and papers
- Cognition 14 launches and papers
- DeepSeek 11 launches and papers
- Meta Superintelligence Labs (MSL) 11 launches and papers
- Mistral AI 11 launches and papers