DeepSeek previews V4 and folds its reasoning line into one million-token model
DeepSeek's first flagship since V3.2 is a single million-token model with selectable thinking depth; the separate R-series ends, and the "R2" the market kept expecting was never released.
- Date
- 24 April 2026
- Who
- DeepSeek
- Confidence
- High on release facts (DeepSeek pages opened); Medium on third-party comparisons
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Landmark · Significance: 4/5 · Org(s): DeepSeek · Confidence: High on release facts (DeepSeek pages opened); Medium on third-party comparisons Primary sources: DeepSeek V4 preview note · DeepSeek updates log · V4-Pro model card · MIT Technology Review · TechCrunch
One-liner. DeepSeek's first flagship since V3.2 is a single million-token model with selectable thinking depth; the separate R-series ends, and the "R2" the market kept expecting was never released.
Why it happened. DeepSeek's line went R1 (reasoning), V3.1 (hybrid, B08-37), V3.2 (agentic reasoning with 10% RL budget, B08-43), and a long-rumoured R2. The Information (relayed by Reuters, 2025-06) reported that Liang Wenfeng was dissatisfied with R2's performance (BGR), and the Financial Times reported in 2025-08 that attempts to train it on Huawei Ascend chips failed, sending DeepSeek back to Nvidia GPUs for training and Ascend for inference (The Register on the FT report); both are reported and DeepSeek has not confirmed either. Distillation and export-control pressure were rising (B08-14).
The idea. V4 is one model family, V4-Pro (1.6T total, 49B active parameters) and V4-Flash (284B total, 13B active), both mixture-of-experts with a 1M-token default context and a hybrid attention design (Compressed Sparse Attention plus Heavily Compressed Attention, building on DSA) that DeepSeek says needs about 27% of V3.2's per-token FLOPs and 10% of its KV cache at 1M tokens for the Pro model, pre-trained on over 32T tokens. Post-training trains domain specialists separately with SFT and RL and then merges them into one model by on-policy distillation. There are three modes, Non-think, Think High and Think Max (API effort low / high / max) (model card, updates log). Weights are MIT-licensed.
Results (company-reported). Codeforces rating 3206, LiveCodeBench 93.5, SWE-bench Verified 80.6, MMLU-Pro 87.5 (V4-Pro); DeepSeek says V4-Pro-Max beats comparable open models and surpasses GPT-5.2 and Gemini 3.0 Pro on some tasks while trailing on knowledge; TechCrunch characterises the models as roughly 3 to 6 months behind the frontier (TechCrunch). Launch API prices were about $1.74 / $3.48 (Pro) and $0.14 / $0.28 (Flash) per million input / output tokens (MIT Technology Review; TechCrunch lists a different Pro input price, so check before quoting). One independent datapoint is that in Deedy Das's July 2026 IMO harness V4 Pro scored 19/42 on a first pass against 42/42 for several rivals (B08-48).
How it spread. V4-Flash had an official release (labelled public beta in the log) on 2026-07-31 and V4-Pro a general-availability release on 2026-08-13; the legacy deepseek-chat and deepseek-reasoner endpoints were discontinued on 2026-07-24; V4.1-Flash with native multimodal support followed on 2026-09-10 (updates log). MIT Technology Review calls V4 "DeepSeek's first model optimized for domestic Chinese chips, such as Huawei's Ascend," mainly for inference, with training possibly still mostly on Nvidia; the "first" is contestable, because the V3.2-Exp model card (2025-09-29, B08-39b) already lists NPU (Ascend) serving images.
Why it mattered. V4 ends "reasoning model" as a separate product class at DeepSeek, as already happened at Anthropic and OpenAI (B08-06), and it shows an efficient-attention route to cheap million-token thinking. The 2025 story of a one-model upset has eased, and coverage now describes the gap to the frontier as months.
Nuance, controversy and myths. "DeepSeek R2" was never released and DeepSeek denied the dated R2 rumours of March 2025 (AIBase report); V4 is not "R2." The 1.6T-parameter figure makes V4-Pro one of the largest open-weight models, and K3 is larger (B08-49). Benchmark claims are DeepSeek's own.
Interview kit.
- 30-second version: V4 is DeepSeek's unified million-token MoE with three thinking depths; it replaced the V-and-R split, reportedly runs on Huawei chips for inference, and coverage puts it a few months behind the frontier.
- Likely follow-ups: What happened to R2? → Never shipped; reportedly delayed over quality and failed Ascend training, then folded into V4 (reported). Is it still cheap? → Yes relative to US frontier prices, but no longer unusual.
- Common mistake: Saying V4 is "R2".
- Connect it to: B08-10, B08-43, B20, B17.
Sources. DeepSeek pages and model card opened; MIT Technology Review and TechCrunch opened; BGR and AIBase (secondary) for R2.