Diffusion language models

Generate blocks of text in parallel by iterative denoising instead of one token at a time, trading some quality for large speed gains.

Autoregressive decoding is sequential and memory-bound. Diffusion LMs refine many tokens per forward pass, reaching 1,000+ tokens/s on one GPU with bidirectional context. The camp sees them as the fast layer for coding and sub-agents now, and possibly a general replacement as quality closes.

Where it stands. Diffusion LMs hold a production niche and are not yet the default. They serve latency-critical coding and sub-agent calls at ~1,000-1,500 tokens/s, while the best-quality models remain autoregressive and the trend is hybrid AR-diffusion decoding.

Evidence

For

  • LLaDA (2025-02-14): an 8B masked-diffusion model trained from scratch is competitive with Llama 3 8B in in-context learning and beats GPT-4o on a reversal-poem task.
  • Mercury Coder (2025-06-17): 1,109 tokens/s (Mini) and 737 (Small) on H100s per Artificial Analysis, up to 10x faster than speed-optimized frontier models at comparable quality.
  • DiffusionGemma (report 2026-07-31): fine-tuned from Gemma 4 MoE (3.8B active, 25.2B total) with under 10% of its training tokens; ~20 tokens per forward pass, ~1,500 tokens/s on one H100.
  • Mercury 2 (reported 2026-06-21): ~1,000 tokens/s and 90 on AIME 2026; Augment Code reports 82% lower latency and 90% lower cost than Claude Opus 4.7 on a compaction sub-agent.
  • Nemotron-Labs-Diffusion (2026-07-07): a tri-mode model where diffusion drafts and autoregression verifies beats multi-token-prediction speculation.

Against

  • DiffusionGemma trails its autoregressive parent: 69.1 vs 88.3 on AIME 2026, and Google's guide recommends Gemma 4 when maximum quality matters (Decrypt, 2026-06-21).
  • Gemini Diffusion matched Gemini 2.0 Flash-Lite on code but lost on science (GPQA Diamond 40.4 vs 56.5) and remains an experimental demo.
  • Mercury 2 is a closed API; local runtimes and agent frameworks are still catching up (Decrypt).
  • Not found: a frontier flagship (GPT-6 Astra, Claude, Gemini 4 Argon) built on diffusion; DiffusionGemma and Nemotron keep AR modes, suggesting hybrids rather than replacement.

Milestones

Who is working on it

  • Stefano Ermon, Aditya Grover, Volodymyr Kuleshov, Inception Labs
  • Shen Nie, Chongxuan Li and co-authors (LLaDA), Renmin University of China
  • Gemini Diffusion and DiffusionGemma teams, Google DeepMind
  • Nemotron-Labs-Diffusion team, NVIDIA

The labs with the most milestones here are Google DeepMind (3), Inception (Inception Labs) (2), NVIDIA (1) and Renmin University of China (1).

Sources

  1. arxiv.org/abs/2502.09992
  2. arxiv.org/abs/2506.17298
  3. deepmind.google/models/gemini-diffusion/
  4. arxiv.org/abs/2608.00146
  5. decrypt.co/371722/inception-labs-mercury-2-ai-beats-googles-diffusiongemma
  6. arxiv.org/abs/2607.05722
  7. blog.google/technology/google-deepmind/gemini-diffusion/
  8. mayfield.com/introducing-inception/

This research bet was checked and corrected against its sources on 6 October 2026. How we check