DeepSeek released R1 and R1-Zero, reasoning models trained with reinforcement learning, with MIT-licensed weights
DeepSeek showed that outcome-checked reinforcement learning on a strong base model produces long, self-correcting chains of thought without human reasoning examples, then released the weights under…
- Date
- 20 January 2025
- Who
- DeepSeek
- People
- Peiyi Wang and Daya Guo (R1-Zero), Junxiao Song (GRPO), Zhihong Shao (first author of DeepSeekMath), Zhibin Gou (distillation), Liang Wenfeng
- Confidence
- High
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Landmark · Significance: 5/5 · Org(s): DeepSeek · People: Peiyi Wang and Daya Guo (R1-Zero), Junxiao Song (GRPO), Zhihong Shao (first author of DeepSeekMath), Zhibin Gou (distillation), Liang Wenfeng · Confidence: High Primary sources: DeepSeek-R1, arXiv 2501.12948 (v1 2025-01-22; v2 2026-01-04) · release post · Nature, 2025-09-17 · DeepSeekMath, arXiv 2402.03300
One-liner. DeepSeek showed that outcome-checked reinforcement learning on a strong base model produces long, self-correcting chains of thought without human reasoning examples, then released the weights under the MIT licence.
Why it happened. OpenAI's o1 proved the capability but published no method. Many open attempts guessed at search, process reward models (PRMs) and MCTS (B07). DeepSeek had two enabling pieces in hand. One was the strong V3-Base, and the other was Group Relative Policy Optimization (GRPO), introduced in the DeepSeekMath paper (arXiv 2024-02-05) as a PPO variant that drops the separate value network and judges each answer against a group of sibling answers (arXiv). The R1 paper's contribution list credits Peiyi Wang and Daya Guo with jointly showing that outcome-based RL induces long chain-of-thought behaviour (creating R1-Zero), and Junxiao Song with proposing the GRPO algorithm, implementing its first version and introducing rule-based math rewards (the contribution note's wording; GRPO's first public description is DeepSeekMath, 2024-02-05, of which Song is a co-author, so for interviews say "GRPO comes from DeepSeekMath, and the R1 team credits Song for the implementation and the rule-based rewards"); Peiyi Wang and Runxin Xu later refined GRPO, Zhibin Gou proposed the large clipping range, and Gou led the distilled series (paper §7, v2). The conceptual bet was that human-written reasoning traces cap performance and that a verifier plus compute is enough.
The idea. Give the base model hard problems whose answers can be checked automatically, let it sample many attempts, reward the correct ones, and the model discovers for itself that thinking longer, checking and backtracking pay off. How it works:
- GRPO: sample 16 answers per question, score each, compute each answer's advantage as (reward minus group mean) divided by group standard deviation, and update with a PPO-style clipped objective (Proximal Policy Optimization, the standard RLHF algorithm) plus a small KL penalty (a term that keeps the policy close to a reference model). No critic network.
- Rewards: rule-based only, with an accuracy reward (boxed final answers matched to ground truth; code run against test cases) and a format reward for
<think>...</think>tags. DeepSeek deliberately avoided neural reward models for reasoning because they are prone to reward hacking at scale. - R1-Zero: RL applied directly to V3-Base with no supervised fine-tuning. The v2 paper reports 10,400 steps, 32 questions x 16 samples per step, maximum length 32,768 tokens and 65,536 after step 8.2K. AIME 2024 pass@1 (accuracy of a single sampled answer) rose from 15.6% to 77.9% (v1 reported 71.0%; with majority voting 86.7%), and average response length grew steadily, with a visible jump in the word "wait," which the authors call the "aha moment."
- R1 (the product): (1) thousands of cold-start long-CoT examples, (2) reasoning RL with an added language-consistency reward to curb Chinese-English mixing, (3) rejection sampling from that checkpoint for about 600K reasoning samples plus about 200K non-reasoning ones (800K total) and two epochs of supervised fine-tuning, (4) a final RL stage mixing rule-based and preference rewards.
- Distillation: six dense models (1.5B to 70B, on Qwen2.5 and Llama bases) fine-tuned only on the 800K samples, no RL. The 32B distillation scored 72.6% on AIME 2024; running the same RL recipe directly on Qwen-32B reached only 47.0%, which the authors read as distillation beating small-model RL (v1 §4.1).
- Reported dead ends: PRMs (hard to define a step, hard to label, reward hacking) and MCTS (token space is too large, value models are hard to train).
Results (v1; company-reported). DeepSeek's abstract says R1 is comparable to OpenAI's o1-1217; the v1 comparison table (Table 4) shows a split. R1 was ahead or level on AIME 2024 (79.8 vs 79.2), MATH-500 (97.3 vs 96.4), LiveCodeBench (65.9 vs 63.4), SWE-bench Verified (49.2 vs 48.9) and DROP (92.2 vs 90.2); o1-1217 was ahead on MMLU (91.8 vs 90.8), GPQA Diamond (75.7 vs 71.5), SimpleQA (47.0 vs 30.1), Codeforces rating (2061 vs 2029) and Aider-Polyglot (61.7 vs 53.3). DeepSeek took the o1 numbers from OpenAI's reports because the API was hard to reach from mainland China. List pricing at launch was $0.55 / $2.19 per million input / output tokens against o1's $15 / $60, about 27x cheaper (release post).
How it spread.
| Lab / project | Response | Date | Lag vs 2025-01-20 |
|---|---|---|---|
| Moonshot AI | Kimi k1.5 (own RL recipe, no weights), announced same day | 2025-01-20 | 0 |
| UC Berkeley (Jiayi Pan et al.) | TinyZero, R1-Zero-style RL on Countdown with a 3B model, "under $30" | about 2025-01-24 | about +4 |
| Hugging Face | Open-R1 plan to rebuild data and training code (blog) | 2025-01-28 | +8 |
| OpenAI | o3-mini, which had been pre-announced on 2024-12-20 for "the end of January," so not counted as an R1 follower; the more visible summaries on 02-06 are plausibly R1-linked (Inference; B08-16) | 2025-01-31 / 02-06 | +11 / +17 |
| Stanford / UW | s1 with 1,000 examples plus "budget forcing" | 2025-01-31 | +11 |
| xAI | Grok 3 Think | 2025-02-17 | +28 |
| Anthropic | Claude 3.7 Sonnet | 2025-02-24 | +35 |
| Alibaba | QwQ-32B via scaled RL, then Qwen3 | 2025-03-06 / 04-29 | +45 / +99 |
| Baidu | ERNIE X1 reasoning model (B08-20a) | 2025-03-16 | +55 |
| NVIDIA | Llama-3.3-Nemotron-Super-49B-v1 with a reasoning on/off prompt (release date; the technical report followed on 2025-05-02) | 2025-03-18 | +57 |
| ByteDance Seed / Tsinghua | DAPO open RL system | 2025-03-18 | +57 |
| Gemini 2.5 Pro | 2025-03-25 | +64 | |
| Mistral | Magistral | 2025-06-10 | +141 |
| MiniMax | MiniMax-M1, open weights, new CISPO RL algorithm | 2025-06-16 | +147 |
| Meta | MobileLLM-R1, small SFT-only reasoning models | 2025-09-12 | +235 |
Diffusion was helped by a short, readable recipe, MIT-licensed weights (including R1-Zero) with distilled small models of 1.5B to 70B parameters, a visible chain of thought, and a price a fraction of the incumbent's. DeepSeek released no training data or RL code, which is why Open-R1, DAPO and others had to reconstruct them.
Why it mattered. It proved (1) reasoning could be reproduced without a secret ingredient, (2) RL with verifiable rewards ("RLVR," named earlier in Tülu 3) is the engine, and (3) a lab then funded by its parent hedge fund, with restricted chips, could ship at the frontier (outside funding for DeepSeek was reported in 2026; see the Backlog). It also redefined open weights as a frontier-competitive category (B19, B20) and set off the market reaction in B08-12.
Nuance, controversy and myths.
- R1 sits on V3-Base. Its extra cost was about $294K (see B08-15); the base model was much more.
- The "aha moment" is a notable observation in one intermediate checkpoint, not proof of understanding; later work questions how much RL creates versus surfaces (B08-25).
- The description of R1-Zero as "no human data" is wrong, because the base model's web pre-training contains vast amounts of reasoning text, which the authors themselves note.
- The v1 and v2 papers report different R1-Zero AIME numbers (71.0% versus 77.9%), so cite the version.
- Distillation from OpenAI was alleged but not publicly proven (B08-14).
Interview kit.
- 30-second version: DeepSeek took a good base model, gave it maths and code problems with automatic answer-checkers, and trained it with a simple RL algorithm (GRPO) until it learned to think at length and self-correct; then it open-sourced everything but the data.
- Likely follow-ups: What is GRPO? → PPO without a value network; advantage is relative to a group of samples. Why no reward model? → Rule-based rewards are hard to hack. Is R1 as good as o1? → DeepSeek says comparable; in its v1 table R1 led on AIME, MATH-500, LiveCodeBench and SWE-bench Verified and trailed o1-1217 on MMLU, GPQA, SimpleQA, Codeforces rating and Aider-Polyglot. Why was it a shock? → Frontier-level reasoning, weights, MIT licence and a very low price at once.
- Common mistake: Saying "R1 proved you don't need much compute." It proved you don't need secret methods; the base model and RL still needed real compute.
- Connect it to: B07, B08-09, B20, B24.
Sources. R1 v2 PDF and v1 PDF (both opened and read locally); release post; DeepSeekMath; Open-R1; TinyZero repo; Nature page redirected to a login hop, so Nature facts come via Scientific American.