Open weights
Models whose weights are published for anyone to download and run. The atlas has logged 368 launches and papers on it since 10 February 2023, 96 of them in 2026. Of those, 107 come with a restricted license.
Alibaba Qwen (Tongyi), DeepSeek and Google DeepMind have the most.
Key launches
The 60 most significant of 368, newest first. Every launch in the atlas
- MiMo-V2.6-Pro and V2.6-Flash Xiaomi (MiMo) · 21 September 2026
Omni-modal 1.02T (Pro) and 310B (Flash) MIT-licensed models; Artificial Analysis rated Pro 46, top of open-weight models at launch. - DeepSeek-V4.1-Flash DeepSeek · 10 September 2026
Native-multimodal 552B-backbone model with a causal encoder-decoder (8B active in prefill, 16B in decode) and a KV cache of 890 bytes per token. - Hy4 preview Tencent · 28 August 2026
770B-total, 49B-active open MoE with 1M-token context, Gated DeepSeek Sparse Attention and hyper-connections, under Apache-2.0. - Qwen3.8-Flash-Next Alibaba (Qwen) · 26 August 2026
125B-parameter (6B active) multimodal MoE that previews the Qwen4 architecture, built on Gated DeltaNet plus Qwen Sparse Attention, with 51B of N-gram embeddings. - Qwen3.8-2.4T-A95B (open weights) Alibaba (Qwen) · 12 August 2026
Open weights of a 2.4T-parameter, 95B-active MoE: Terminal Bench 2.1 86.6 and SWE-bench Pro 67.7, under a Qwen3.8-Max license. - Muse Glimmer 30B Meta · 10 August 2026
Meta's first open-weights model since Llama 4: 30B dense, Apache 2.0, logit-distilled from Muse Spark, built to run locally. - MiniMax H3 (Hailuo) MiniMax · 31 July 2026
MiniMax H3 is an open-weights audio-video model that makes 2K, 15-second clips with native stereo sound, priced below a third of mainstream models at 2K. - Kimi K3 Moonshot AI · 16 July 2026
2.8T-parameter open-weights MoE (104B active, 1M context) on Kimi Delta Attention and Attention Residuals, reported third on Artificial Analysis behind two closed models at launch. - Hy3 Tencent · 6 July 2026
Completed Hy3: same 295B/21B MoE, now Apache-2.0, API input at 1 yuan per million tokens; scored 2.67 vs GLM-5.1's 2.51 in a 270-expert blind test. - LongCat-2.0 Meituan (LongCat) · 29 June 2026
1.6T-total MoE with 1M context, trained on 35T+ tokens entirely on Chinese AI accelerators; 59.5 on SWE-bench Pro, weights under MIT. - GLM-5.2 Zhipu AI / Z.ai · 13 June 2026
Z.ai's flagship GLM with, for the first time, a solid 1M-token context; 62.1 on SWE-Bench Pro and 40.5 on HLE, MIT-licensed, weights following three days after subscriber launch. - MiniMax M3 MiniMax · 1 June 2026
Natively multimodal ~428B-total (~23B active) model with MiniMax Sparse Attention for 1M context at about 1/20 the per-token cost of M2; 80.5% SWE-bench Verified. - DeepSeek-V4-Pro and V4-Flash (preview) DeepSeek · 24 April 2026
1.6T-parameter MoE (49B active) plus 284B Flash, 1M-token context; at 1M tokens uses 27% of V3.2's FLOPs and 10% of its KV cache. - DeepSeek-V4 (CSA/HCA attention, mHC, Muon) DeepSeek · 24 April 2026
1.6T-parameter V4-Pro (49B active) and 284B V4-Flash (13B active) with 1M-token context; Pro uses 27% of V3.2's inference FLOPs and 10% of its KV cache. - MiMo-V2.5 and V2.5-Pro Xiaomi (MiMo) · 22 April 2026
Xiaomi opened a 1.02T-total (42B active) Pro model and a 310B/15B model under MIT, with 1M context. - GLM-5.1 Zhipu AI / Z.ai · 7 April 2026
Post-trained GLM-5 that tops SWE-Bench Pro at 58.4 and can work autonomously on one task for up to eight hours. - Gemma 4 (E2B, E4B, 26B MoE, 31B) Google DeepMind · 2 April 2026
Gemma 4 ships under Apache 2.0 in four sizes; the 31B dense ranked 3 among open models on Arena AI text, the 26B MoE 6. - Qwen3.5-397B-A17B and Qwen3.5-Plus Alibaba (Qwen) · 16 February 2026
Open 397B-A17B native vision-language MoE with Gated DeltaNet linear attention, 201 languages and 262K context; hosted Qwen3.5-Plus offers 1M. - GLM-5 Zhipu AI / Z.ai · 11 February 2026
744B-total (40B active) MIT-licensed MoE with DeepSeek Sparse Attention, trained on 28.5T tokens; 77.8% SWE-bench Verified and top open model on Artificial Analysis at launch. - Kimi K2.5 Moonshot AI · 27 January 2026
Natively multimodal 1T/32B open model trained on about 15T mixed vision-text tokens, with Agent Swarm orchestrating up to 100 parallel sub-agents. - DeepSeek-V3.2 (DSA, scaled RL, agentic synthesis) DeepSeek · 2 December 2025
Pairs sparse attention with a scaled RL budget and a synthetic agentic-task pipeline; V3.2-Speciale claims IMO and IOI 2025 gold-level results. - Mistral 3: Mistral Large 3 and Ministral 3 Mistral AI · 2 December 2025
Mistral Large 3, a 675B-parameter MoE with 41B active trained on 3,000 H200s, and Ministral 3 (3B/8B/14B) all ship under Apache 2.0. - DeepSeek-V3.2 and V3.2-Speciale DeepSeek · 1 December 2025
Production DSA model at claimed GPT-5 level; Speciale variant claims gold at IMO, CMO, ICPC World Finals and IOI 2025. First thinking-in-tool-use release. - FLUX.2 [pro / flex / dev / klein] Black Forest Labs · 25 November 2025
FLUX.2 pairs a Mistral-3 24B vision-language model with a rectified-flow transformer, and supports up to 10 reference images and 4MP editing, with a 32B open [dev] model. - OpenClaw (Clawdbot / Moltbot) OpenClaw (Peter Steinberger) · 24 November 2025
Peter Steinberger releases an open-source, self-hosted personal agent (first Clawdbot, then Moltbot, then OpenClaw) driven from WhatsApp, Telegram and Discord. - SAM 3 and SAM 3D Meta · 19 November 2025
SAM 3 segments and tracks every instance matching a text or exemplar concept; SAM 3D reconstructs objects and human bodies from one image. - Kimi K2 Thinking Moonshot AI · 6 November 2025
Open-weights thinking agent that interleaves reasoning with 200-300 sequential tool calls; claims state of the art on HLE with tools (44.9%) and BrowseComp (60.2%). - Getty Images v Stability AI (UK High Court judgment) Getty Images / Stability AI · 4 November 2025
UK High Court rejected Getty's secondary copyright claim, holding that Stable Diffusion's weights are not an 'infringing copy', and made only narrow trade-mark findings for… - Kimi Linear Moonshot AI · 30 October 2025
Hybrid linear-attention architecture (Kimi Delta Attention + MLA) that beats full attention in fair comparisons while cutting KV cache up to 75%. - MiniMax-M2 MiniMax · 27 October 2025
230B-total, 10B-active open MoE built for coding and agents with interleaved thinking; scored 61 on the Artificial Analysis index, ranked first among open models. - DeepSeek-V3.2-Exp (DeepSeek Sparse Attention) DeepSeek · 29 September 2025
First production use of DeepSeek Sparse Attention, with V3.1-Terminus-level quality, lower long-context cost and API prices cut by over 50%. - DeepSeek Sparse Attention (V3.2-Exp) DeepSeek · 29 September 2025
DeepSeek Sparse Attention (DSA) in V3.2-Exp is fine-grained sparse attention with output quality near V3.1-Terminus, which enabled a 50%+ API price cut. - Qwen3-VL (235B-A22B first; 2B to 32B and 30B-A3B in October) Alibaba (Qwen) · 23 September 2025
Open 235B-A22B vision-language flagship with 256K context (1M extendable), OCR in 32 languages and GUI-agent skills. Smaller sizes followed in October. - DeepSeek-R1 in Nature (peer-reviewed) DeepSeek · 17 September 2025
R1 becomes the first major LLM paper through peer review; supplement discloses roughly $294K RL cost on top of the base model. - Qwen3-Next-80B-A3B (Instruct and Thinking) Alibaba (Qwen) · 11 September 2025
80B MoE with 3B active, mixing Gated DeltaNet and gated attention 3:1; claimed 10x throughput beyond 32K context at 10% of Qwen3-32B training cost. - DeepSeek-V3.1 DeepSeek · 21 August 2025
One 671B model with switchable thinking and non-thinking modes, stronger tool use and a UE8M0 FP8 format; DeepSeek's stated 'first step toward the agent era'. - gpt-oss-120b and gpt-oss-20b OpenAI · 5 August 2025
OpenAI's first open-weight language models since GPT-2: Apache 2.0 reasoning models, 120B near o4-mini and 20B near o3-mini. - Qwen-Image (20B MMDiT) Alibaba (Qwen) · 4 August 2025
20B MMDiT image model with native text rendering for English and Chinese, opened under Apache 2.0 with a technical report. - Qwen-Image Alibaba (Qwen) · 4 August 2025
Qwen-Image is a 20B MMDiT image foundation model, released with code and weights, that claims leading results on complex text rendering, especially Chinese. - GLM-4.5 and GLM-4.5-Air Zhipu AI / Z.ai · 28 July 2025
355B-total (32B active) open MoE that unifies reasoning, coding and agent tool use in hybrid thinking and non-thinking modes, MIT licensed. - Wan2.2 (T2V-A14B, I2V-A14B, TI2V-5B) Alibaba (Qwen) · 28 July 2025
Open video diffusion adopts mixture-of-experts. The A14B experts split denoising by timestep, and a 5B hybrid model gives 720P 24fps on a 4090. - Kimi K2 technical report (MuonClip) Moonshot AI · 28 July 2025
1T-parameter MoE (32B active) pretrained on 15.5T tokens with no loss spikes using MuonClip, plus a large-scale agentic data synthesis pipeline. - Qwen3-Coder-480B-A35B and Qwen Code Alibaba (Qwen) · 22 July 2025
480B-A35B open agentic coding model, 256K native context (1M with YaRN), claimed best open model on agentic coding and competitive with Claude Sonnet 4. - Kimi K2 Moonshot AI · 11 July 2025
1T-parameter (32B active) open-weights MoE trained with the MuonClip optimizer on 15.5T tokens, with zero reported loss spikes and a focus on agentic tool use. - ERNIE 4.5 open-source family Baidu · 30 June 2025
Baidu open-sourced ten ERNIE 4.5 models under Apache 2.0, from a 0.3B dense model to a 424B-total (47B active) multimodal MoE. - Gemini CLI Google · 25 June 2025
Gemini CLI: Apache-2.0 terminal coding agent with free Gemini 2.5 Pro access, 60 requests per minute and 1,000 per day on a personal Google account. - MiniMax-M1 MiniMax · 16 June 2025
Open 456B hybrid-attention reasoning model with 1M-token context and the CISPO RL algorithm; full RL run cost a reported $534,700. - V-JEPA 2 Meta · 11 June 2025
1.2B-param video world model pretrained on 1M+ hours of video, then adapted on only 62 hours of robot data for zero-shot pick-and-place planning. - Magistral Small and Medium Mistral AI · 10 June 2025
Magistral is Mistral's first reasoning model. Medium scores 73.6% on AIME 2024 (90% with majority voting), and Small is open under Apache 2.0. - Qwen3 (0.6B to 235B-A22B) Alibaba (Qwen) · 29 April 2025
Eight open Apache 2.0 models, two MoE (235B-A22B, 30B-A3B) and six dense, with switchable thinking modes, 36T tokens and 119 languages. - Llama 4 Scout and Maverick (Behemoth previewed) Meta · 5 April 2025
The first MoE Llamas are Scout (17B active/109B total, 10M context) and Maverick (17B/400B, 128 experts), while the 2T-param Behemoth was only previewed. - DAPO ByteDance Seed / Tsinghua AIR · 18 March 2025
Fully open large-scale RL system reaches 50 on AIME 2024 with Qwen2.5-32B, beating R1-Zero-Qwen-32B in half the training steps. - DeepSeek-R1: Incentivizing Reasoning via RL (arXiv) DeepSeek · 22 January 2025
Shows reasoning can emerge from pure RL on verifiable rewards (R1-Zero), then builds R1 with a small cold-start and distills it to small models. - DeepSeek-R1 and R1-Zero DeepSeek · 20 January 2025
Open-weights 671B MoE reasoning model claimed at OpenAI o1 level; R1-Zero showed reasoning emerging from pure RL with rule-based rewards. - DeepSeek-V3 Technical Report DeepSeek · 27 December 2024
671B MoE (37B active) trained on 14.8T tokens with 2.788M H800 GPU-hours, no loss spikes or rollbacks; auxiliary-loss-free balancing and multi-token prediction. - DeepSeek-V3 DeepSeek · 26 December 2024
671B MoE (37B active) trained on 14.8T tokens in 2.788M H800 GPU hours; open weights matching leading closed models at far lower reported cost. - Llama 3.1 (8B, 70B, 405B) Meta · 23 July 2024
Llama 3.1 405B: first open-weights model Meta says is competitive with GPT-4-class closed models, 128K context. - DeepSeekMath (introduces GRPO) DeepSeek · 5 February 2024
Introduces Group Relative Policy Optimization (GRPO), a critic-free PPO variant, plus a 7B math model scoring 51.7% on MATH. - Llama 2 Meta · 18 July 2023
Llama 2 7B-70B plus Llama 2-Chat released free for research and commercial use, with Microsoft as preferred cloud partner. - LLaMA Meta · 24 February 2023
Meta's 7B-65B LLaMA trained on public data only; 13B beats GPT-3 175B on most benchmarks. Weights leaked on 4chan within a week.
Most active labs
- Alibaba Qwen (Tongyi) 64 launches and papers
- DeepSeek 53 launches and papers
- Google DeepMind 27 launches and papers
- Mistral AI 27 launches and papers
- Z.ai (Zhipu AI) 18 launches and papers
- Meta 16 launches and papers
- Moonshot AI (Kimi) 13 launches and papers
- Microsoft 12 launches and papers