Diffusion language models
Generate blocks of text in parallel by iterative denoising instead of one token at a time, trading some quality for large speed gains.
Autoregressive decoding is sequential and memory-bound. Diffusion LMs refine many tokens per forward pass, reaching 1,000+ tokens/s on one GPU with bidirectional context. The camp sees them as the fast layer for coding and sub-agents now, and possibly a general replacement as quality closes.
Where it stands. Diffusion LMs hold a production niche and are not yet the default. They serve latency-critical coding and sub-agent calls at ~1,000-1,500 tokens/s, while the best-quality models remain autoregressive and the trend is hybrid AR-diffusion decoding.
Evidence
For
- LLaDA (2025-02-14): an 8B masked-diffusion model trained from scratch is competitive with Llama 3 8B in in-context learning and beats GPT-4o on a reversal-poem task.
- Mercury Coder (2025-06-17): 1,109 tokens/s (Mini) and 737 (Small) on H100s per Artificial Analysis, up to 10x faster than speed-optimized frontier models at comparable quality.
- DiffusionGemma (report 2026-07-31): fine-tuned from Gemma 4 MoE (3.8B active, 25.2B total) with under 10% of its training tokens; ~20 tokens per forward pass, ~1,500 tokens/s on one H100.
- Mercury 2 (reported 2026-06-21): ~1,000 tokens/s and 90 on AIME 2026; Augment Code reports 82% lower latency and 90% lower cost than Claude Opus 4.7 on a compaction sub-agent.
- Nemotron-Labs-Diffusion (2026-07-07): a tri-mode model where diffusion drafts and autoregression verifies beats multi-token-prediction speculation.
Against
- DiffusionGemma trails its autoregressive parent: 69.1 vs 88.3 on AIME 2026, and Google's guide recommends Gemma 4 when maximum quality matters (Decrypt, 2026-06-21).
- Gemini Diffusion matched Gemini 2.0 Flash-Lite on code but lost on science (GPQA Diamond 40.4 vs 56.5) and remains an experimental demo.
- Mercury 2 is a closed API; local runtimes and agent frameworks are still catching up (Decrypt).
- Not found: a frontier flagship (GPT-6 Astra, Claude, Gemini 4 Argon) built on diffusion; DiffusionGemma and Nemotron keep AR modes, suggesting hybrids rather than replacement.
Milestones
- DiffusionGemma technical report describes an open-weight diffusion LM fine-tuned from Gemma 4 Google DeepMind · 31 July 2026
- Nemotron-Labs-Diffusion unifies autoregressive, diffusion and self-speculation decoding NVIDIA · 7 July 2026
- Decrypt compares Inception's Mercury 2 with Google's DiffusionGemma Inception Labs / Google DeepMind · 21 June 2026
- Mercury, commercial diffusion LLMs for code Inception Labs · 17 June 2025
- Gemini Diffusion experimental text-diffusion demo (about 1,479 tokens/s reported) Google DeepMind · May 2025
- LLaDA: Large Language Diffusion Models, 8B trained from scratch (arXiv v1) Renmin University of China · 14 February 2025
Who is working on it
- Stefano Ermon, Aditya Grover, Volodymyr Kuleshov, Inception Labs
- Shen Nie, Chongxuan Li and co-authors (LLaDA), Renmin University of China
- Gemini Diffusion and DiffusionGemma teams, Google DeepMind
- Nemotron-Labs-Diffusion team, NVIDIA
The labs with the most milestones here are Google DeepMind (3), Inception (Inception Labs) (2), NVIDIA (1) and Renmin University of China (1).
Sources
- arxiv.org/abs/2502.09992
- arxiv.org/abs/2506.17298
- deepmind.google/models/gemini-diffusion/
- arxiv.org/abs/2608.00146
- decrypt.co/371722/inception-labs-mercury-2-ai-beats-googles-diffusiongemma
- arxiv.org/abs/2607.05722
- blog.google/technology/google-deepmind/gemini-diffusion/
- mayfield.com/introducing-inception/
This research bet was checked and corrected against its sources on 6 October 2026. How we check