DeepSeek-V3.2 and V3.2-Speciale spend more than 10% of pre-training cost on RL

Ten months after R1's cheap RL run, DeepSeek reported a stable RL protocol that consumed more than a tenth of pre-training compute, plus an open model with gold-level contest results.

Date
1 December 2025
Who
DeepSeek
Confidence
High
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Landmark · Significance: 4/5 · Org(s): DeepSeek · Confidence: High Primary sources: V3.2 paper, arXiv 2512.02556 (v1 2025-12-02; paper read locally) · DeepSeekMath-V2, arXiv 2511.22570 (2025-11-27)

One-liner. Ten months after R1's cheap RL run, DeepSeek reported a stable RL protocol that consumed more than a tenth of pre-training compute, plus an open model with gold-level contest results.

Why it happened. R1 proved the recipe at about $294K of RL. Through 2025 closed labs reported ever larger RL runs (o3 at a reported 10x o1's training compute, read by Epoch as RL compute; Grok 4's "pretraining-scale" RL), and DeepSeek's models lagged on agentic and software-engineering evaluations (B08-12).

The idea. Three changes. (1) DeepSeek Sparse Attention (DSA), an attention mechanism that cuts long-context compute. (2) A "stable and scalable" RL protocol that allocates a post-training budget "exceeding 10% of the pre-training cost" and shows performance rising with RL budget, with the authors hypothesising further gains from more. (3) A synthesis pipeline for agentic tasks that integrates reasoning with tool use in single trajectories. A high-compute variant, V3.2-Speciale, was trained with relaxed length constraints (paper).

Results (company-reported). V3.2 performs similarly to GPT-5-high on reasoning tasks and slightly worse than Gemini 3.0 Pro. Speciale scored IMO 2025 35/42, CMO 2025 102/126, IOI 2025 492/600 (10th) and ICPC World Finals 2025 10/12 (second), all at gold level, with weaker token efficiency than Gemini 3 Pro (the authors say they imposed stricter token limits on the official V3.2 to control cost). DeepSeekMath-V2 (2025-11-27) targeted a gap that answer-only rewards leave. It trains an LLM verifier that scores proofs for rigour and uses it as the reward for a proof generator, and reports gold-level IMO 2025, gold-level CMO 2024 and 118/120 on Putnam 2024 with scaled test-time compute (arXiv 2511.22570).

How it spread. The open-weight frontier moved from imitation to scaling. Moonshot shipped open agentic-reasoning models (K2 Thinking, K3) and, per secondary reports, Qwen, GLM and MiniMax followed through 2026 (B08-51). Methodologically, the verifier-as-reward idea answers the Yue et al. critique that answer-only RL cannot expand capability, and it anticipates the proof-level IMO 2026 results (B08-48).

Why it mattered. It documents the end of the "reasoning is nearly free" window, because the RL stage became a major line item and the open-weights gap to closed labs depends on RL compute as well as architecture. It was also the point where DeepSeek's reasoning and tool-use work merged in one model (B08-47).

Nuance, controversy and myths. The gold-medal claims come from DeepSeek's own evaluation protocol (the paper's Appendix D), and the contests did not grade them officially. Speciale is a research-grade, token-hungry variant and is not the default API model. "More than 10% of pre-training" is a lower bound DeepSeek chose to state and leaves out the full cost.

Interview kit.

  • 30-second version: In late 2025 DeepSeek reported that its RL post-training now cost over 10% of pre-training and that performance kept scaling with it; its high-compute Speciale variant matched gold-level scores at IMO, IOI and the ICPC finals.
  • Likely follow-ups: How does that square with "$294K"? → R1's RL was tiny; scaling it up is the point. Why a verifier for proofs? → Right answers do not imply right reasoning (DeepSeekMath-V2).
  • Common mistake: Citing the V3.2 results as independently graded.
  • Connect it to: B08-15, B08-40, B08-34, B02, B16.

Sources. V3.2 paper, DeepSeekMath-V2 abstract and the DeepSeek release note opened; the arXiv listing dates 2025-12-02, the product announcement 2025-12-01.

Read it in the deep dive