DeepSeek-V3.2-Exp adds sparse attention and cuts API prices

DeepSeek released V3.2-Exp on 2025-09-29 (252 days after R1), an experimental model that introduced DeepSeek Sparse Attention (DSA), a fine-grained sparse attention for faster training and inference…

Date
29 September 2025
Who
DeepSeek
Confidence
High (DeepSeek's release note and model card opened)
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): DeepSeek · Confidence: High (DeepSeek's release note and model card opened) DeepSeek released V3.2-Exp on 2025-09-29 (252 days after R1), an experimental model that introduced DeepSeek Sparse Attention (DSA), a fine-grained sparse attention for faster training and inference on long contexts, with performance on par with V3.1-Terminus; it cut API prices by more than half the same day, open-sourced the weights, a technical report and GPU kernels (TileLang and CUDA), and kept V3.1-Terminus on the API until 2025-10-15 for side-by-side comparison (DeepSeek release note). The model card lists SGLang container images for NPU (Ascend) serving next to GPU ones, which matters for the later "first Ascend-optimised model" claim about V4 (B08-47; model card). DSA is the efficiency base of V3.2 and, built upon, of V4, and the price cut is an early example of the RL-era price war (B08-06, B18). Sources: release note · model card

Read it in the deep dive