Efficient architectures: hybrid, linear and sparse attention

Replace quadratic full attention with linear, recurrent or sparse layers so million-token contexts and long agent loops stay affordable.

Long-context and agentic workloads make inference cost, not training cost, the constraint. Hybrids (Gated DeltaNet or Mamba layers with a few attention layers) and learned sparse attention cut KV cache and decode cost several-fold while keeping quality, so labs from Alibaba to DeepSeek are adopting them.

Where it stands. Hybrid and sparse attention is now common among open-weight labs (Qwen, DeepSeek, Moonshot, MiniMax) and NVIDIA/IBM; closed frontier labs disclose nothing, so adoption at the very top is unknown (not found).

Evidence

For

  • Qwen3-Next-80B-A3B (Sep 2025): Gated DeltaNet plus gated attention hybrid; base beats Qwen3-32B-Base on ~10% of the training cost with ~10x throughput beyond 32K tokens (company-claimed).
  • Kimi Linear (2025-10-30): a 3:1 KDA-to-MLA hybrid beat full attention in fair comparisons, cutting KV cache up to 75% and giving up to 6x decode throughput at 1M tokens.
  • Qwen3.5 (2026-02-16) uses 3 Gated DeltaNet layers per attention layer at 397B parameters; Qwen3.8-Flash-Next (Aug 2026) previews the Qwen4 design.
  • DeepSeek's sparse attention arrived with V3.2-Exp (2025-09-29, API prices cut 50%+); V4 (2026-04-24) made 1M context the default; V4.1-Flash (2026-09-10) needs 1/4 the HBM for KV cache.
  • Kimi K3 (2026-07-17): a 2.8T-parameter open model built on Kimi Delta Attention (linear) and Attention Residuals with 1M context; Moonshot reports ~2.5x scaling efficiency over K2 but trails Claude Fable 5 and GPT-5.6 Sol.
  • NVIDIA Nemotron 3 (2025-12-15) and IBM Granite 4.0 (2025-10-02) ship open hybrid Mamba-Transformer models.

Against

  • MiniMax reverted to full attention for M2 (2025-10-29): efficient-attention deficits in multi-hop reasoning appeared only at larger scale, and leaderboards had not shown them.
  • Inference: the best design at frontier scale is unsettled: Moonshot's K3 bets on linear hybrids while MiniMax's M3 (June 2026) chose sparse attention.
  • Mamba-3's authors (2026-03-16) note many linear models trade away quality, for example failing at state tracking.
  • TTT-E2E (2025-12-29) reports Mamba 2 and Gated DeltaNet do not scale with context length like full attention in its 3B-parameter setting.
  • Qwen's own hybrids keep one full-attention layer in four; closed frontier labs do not disclose attention design, so use at the top is not found.

Milestones

Who is working on it

  • Qwen team, Alibaba Qwen
  • DeepSeek team (NSA, DSA, V4 attention), DeepSeek
  • Kimi Team (Kimi Linear), Moonshot AI
  • Albert Gu, Tri Dao, J. Zico Kolter (Mamba-3), CMU / Princeton
  • Nemotron team, NVIDIA
  • Granite team, IBM
  • MiniMax pretraining team, MiniMax

The labs with the most milestones here are DeepSeek (5), Alibaba Qwen (Tongyi) (3), MiniMax (2), Moonshot AI (Kimi) (2), CMU / Princeton (1) and IBM (1).

Sources

  1. huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct
  2. huggingface.co/Qwen/Qwen3.5-397B-A17B
  3. huggingface.co/Qwen/Qwen3.8-Flash-Next
  4. arxiv.org/abs/2510.26692
  5. kimi.com/blog/kimi-k3
  6. api-docs.deepseek.com/news/news250929
  7. api-docs.deepseek.com/news/news260424
  8. api-docs.deepseek.com/news/news260910
  9. minimax.io/news/why-did-m2-end-up-as-a-full-attention-model
  10. arxiv.org/abs/2606.13392
  11. arxiv.org/abs/2603.15569
  12. arxiv.org/abs/2601.07372
  13. arxiv.org/abs/2512.23675
  14. ibm.com/new/announcements/ibm-granite-4-0-hyper-efficient-high-performance-hybrid-models
  15. developer.nvidia.com/blog/inside-nvidia-nemotron-3-techniques-tools-and-data-that-make-it-
  16. minimax.io/blog/minimax-m3
  17. arxiv.org/abs/2502.11089

This research bet was checked against its sources on 6 October 2026. How we check