Efficient architectures: hybrid, linear and sparse attention
Replace quadratic full attention with linear, recurrent or sparse layers so million-token contexts and long agent loops stay affordable.
Long-context and agentic workloads make inference cost, not training cost, the constraint. Hybrids (Gated DeltaNet or Mamba layers with a few attention layers) and learned sparse attention cut KV cache and decode cost several-fold while keeping quality, so labs from Alibaba to DeepSeek are adopting them.
Where it stands. Hybrid and sparse attention is now common among open-weight labs (Qwen, DeepSeek, Moonshot, MiniMax) and NVIDIA/IBM; closed frontier labs disclose nothing, so adoption at the very top is unknown (not found).
Evidence
For
- Qwen3-Next-80B-A3B (Sep 2025): Gated DeltaNet plus gated attention hybrid; base beats Qwen3-32B-Base on ~10% of the training cost with ~10x throughput beyond 32K tokens (company-claimed).
- Kimi Linear (2025-10-30): a 3:1 KDA-to-MLA hybrid beat full attention in fair comparisons, cutting KV cache up to 75% and giving up to 6x decode throughput at 1M tokens.
- Qwen3.5 (2026-02-16) uses 3 Gated DeltaNet layers per attention layer at 397B parameters; Qwen3.8-Flash-Next (Aug 2026) previews the Qwen4 design.
- DeepSeek's sparse attention arrived with V3.2-Exp (2025-09-29, API prices cut 50%+); V4 (2026-04-24) made 1M context the default; V4.1-Flash (2026-09-10) needs 1/4 the HBM for KV cache.
- Kimi K3 (2026-07-17): a 2.8T-parameter open model built on Kimi Delta Attention (linear) and Attention Residuals with 1M context; Moonshot reports ~2.5x scaling efficiency over K2 but trails Claude Fable 5 and GPT-5.6 Sol.
- NVIDIA Nemotron 3 (2025-12-15) and IBM Granite 4.0 (2025-10-02) ship open hybrid Mamba-Transformer models.
Against
- MiniMax reverted to full attention for M2 (2025-10-29): efficient-attention deficits in multi-hop reasoning appeared only at larger scale, and leaderboards had not shown them.
- Inference: the best design at frontier scale is unsettled: Moonshot's K3 bets on linear hybrids while MiniMax's M3 (June 2026) chose sparse attention.
- Mamba-3's authors (2026-03-16) note many linear models trade away quality, for example failing at state tracking.
- TTT-E2E (2025-12-29) reports Mamba 2 and Gated DeltaNet do not scale with context length like full attention in its 3B-parameter setting.
- Qwen's own hybrids keep one full-attention layer in four; closed frontier labs do not disclose attention design, so use at the top is not found.
Milestones
- DeepSeek-V4.1-Flash is a causal encoder-decoder with a KV cache using 1/4 the HBM of the prior generation DeepSeek · 10 September 2026
- Qwen3.8-Flash-Next previews the Qwen4 architecture, which combines Gated DeltaNet with Qwen Sparse Attention Alibaba Qwen · August 2026
- Kimi K3: 2.8T-parameter open model on Kimi Delta Attention and Attention Residuals, 1M context Moonshot AI · 17 July 2026
- MiniMax Sparse Attention paper accompanies MiniMax-M3 (1M context) MiniMax · 11 June 2026
- DeepSeek V4 preview makes 1M context the default via token compression plus sparse attention DeepSeek · 24 April 2026
- Mamba-3: inference-first state space model CMU / Princeton · 16 March 2026
- Qwen3.5-397B-A17B: Gated DeltaNet and MoE hybrid, early-fusion multimodal Alibaba Qwen · 16 February 2026
- DeepSeek Engram adds conditional memory via scalable lookup as a new sparsity axis beside MoE DeepSeek · 12 January 2026
- NVIDIA Nemotron 3: open hybrid Mamba-Transformer MoE family NVIDIA · 15 December 2025
- Kimi Linear, a hybrid linear attention that beats full attention in fair comparisons Moonshot AI · 30 October 2025
- MiniMax explains why M2 returned to full attention MiniMax · 29 October 2025
- IBM Granite 4.0 hybrid Mamba/transformer models (Apache 2.0) IBM · 2 October 2025
- DeepSeek-V3.2-Exp debuts DeepSeek Sparse Attention DeepSeek · 29 September 2025
- Qwen3-Next-80B-A3B released with Gated DeltaNet hybrid attention Alibaba Qwen · September 2025
- Native Sparse Attention, hardware-aligned and natively trainable sparse attention (arXiv v1) DeepSeek · 16 February 2025
Who is working on it
- Qwen team, Alibaba Qwen
- DeepSeek team (NSA, DSA, V4 attention), DeepSeek
- Kimi Team (Kimi Linear), Moonshot AI
- Albert Gu, Tri Dao, J. Zico Kolter (Mamba-3), CMU / Princeton
- Nemotron team, NVIDIA
- Granite team, IBM
- MiniMax pretraining team, MiniMax
The labs with the most milestones here are DeepSeek (5), Alibaba Qwen (Tongyi) (3), MiniMax (2), Moonshot AI (Kimi) (2), CMU / Princeton (1) and IBM (1).
Sources
- huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct
- huggingface.co/Qwen/Qwen3.5-397B-A17B
- huggingface.co/Qwen/Qwen3.8-Flash-Next
- arxiv.org/abs/2510.26692
- kimi.com/blog/kimi-k3
- api-docs.deepseek.com/news/news250929
- api-docs.deepseek.com/news/news260424
- api-docs.deepseek.com/news/news260910
- minimax.io/news/why-did-m2-end-up-as-a-full-attention-model
- arxiv.org/abs/2606.13392
- arxiv.org/abs/2603.15569
- arxiv.org/abs/2601.07372
- arxiv.org/abs/2512.23675
- ibm.com/new/announcements/ibm-granite-4-0-hyper-efficient-high-performance-hybrid-models
- developer.nvidia.com/blog/inside-nvidia-nemotron-3-techniques-tools-and-data-that-make-it-
- minimax.io/blog/minimax-m3
- arxiv.org/abs/2502.11089
This research bet was checked against its sources on 6 October 2026. How we check