DeepSeek-R1's last stage is preference RL

The DeepSeek-R1 report, which anchors the reasoning-RL chapters (B08), also shows that RLHF-style preference RL did not disappear.

Date
22 January 2025
Who
DeepSeek-AI
Confidence
High (the report's own description)
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): DeepSeek-AI · Confidence: High (the report's own description) The DeepSeek-R1 report, which anchors the reasoning-RL chapters (B08), also shows that RLHF-style preference RL did not disappear. After reasoning RL with rule-based rewards and a rejection-sampling SFT round (about 800k samples), section 2.3.4 adds a second RL stage whose stated aim is better alignment with human preferences, meaning more helpfulness and harmlessness, while still refining reasoning. For reasoning prompts it keeps rule-based rewards, but for general data it uses reward models to capture human preferences, scoring helpfulness on the final summary only and harmlessness on the whole response, reasoning included (arXiv:2501.12948, v1 2025-01-22, §2.3.4). So "RLHF versus reasoning RL" is a false split inside one pipeline. Verifiable rewards drive the reasoning, and preference-model rewards shape style and safety. This is the cleanest answer to the myth in the Myths section. Sources: DeepSeek-R1 report

Read it in the deep dive