DeepSeek-V3.2-Exp adds sparse attention and cuts API prices
DeepSeek released V3.2-Exp on 2025-09-29 (252 days after R1), an experimental model that introduced DeepSeek Sparse Attention (DSA), a fine-grained sparse attention for faster training and inference…
- Date
- 29 September 2025
- Who
- DeepSeek
- Confidence
- High (DeepSeek's release note and model card opened)
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): DeepSeek · Confidence: High (DeepSeek's release note and model card opened) DeepSeek released V3.2-Exp on 2025-09-29 (252 days after R1), an experimental model that introduced DeepSeek Sparse Attention (DSA), a fine-grained sparse attention for faster training and inference on long contexts, with performance on par with V3.1-Terminus; it cut API prices by more than half the same day, open-sourced the weights, a technical report and GPU kernels (TileLang and CUDA), and kept V3.1-Terminus on the API until 2025-10-15 for side-by-side comparison (DeepSeek release note). The model card lists SGLang container images for NPU (Ascend) serving next to GPU ones, which matters for the later "first Ascend-optimised model" claim about V4 (B08-47; model card). DSA is the efficiency base of V3.2 and, built upon, of V4, and the price cut is an early example of the RL-era price war (B08-06, B18). Sources: release note · model card