DeepSeek-R1's last stage is preference RL
The DeepSeek-R1 report, which anchors the reasoning-RL chapters (B08), also shows that RLHF-style preference RL did not disappear.
- Date
- 22 January 2025
- Who
- DeepSeek-AI
- Confidence
- High (the report's own description)
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): DeepSeek-AI · Confidence: High (the report's own description) The DeepSeek-R1 report, which anchors the reasoning-RL chapters (B08), also shows that RLHF-style preference RL did not disappear. After reasoning RL with rule-based rewards and a rejection-sampling SFT round (about 800k samples), section 2.3.4 adds a second RL stage whose stated aim is better alignment with human preferences, meaning more helpfulness and harmlessness, while still refining reasoning. For reasoning prompts it keeps rule-based rewards, but for general data it uses reward models to capture human preferences, scoring helpfulness on the final summary only and harmlessness on the whole response, reasoning included (arXiv:2501.12948, v1 2025-01-22, §2.3.4). So "RLHF versus reasoning RL" is a false split inside one pipeline. Verifiable rewards drive the reasoning, and preference-model rewards shape style and safety. This is the cleanest answer to the myth in the Myths section. Sources: DeepSeek-R1 report