DAPO and Dr. GRPO publish open RL recipes that R1 left out

DAPO ("Decoupled Clip and Dynamic sAmpling Policy Optimization," arXiv v1 2025-03-18) states in its abstract that key technical details of top reasoning LLMs are concealed in both the o1 blog and the…

Date
18 March 2025
Who
ByteDance Seed, Tsinghua AIR; Sea AI Lab
People
Qiying Yu et al.; Zichen Liu, Min Lin et al.
Confidence
High
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 4/5 · Org(s): ByteDance Seed, Tsinghua AIR; Sea AI Lab · People: Qiying Yu et al.; Zichen Liu, Min Lin et al. · Confidence: High DAPO ("Decoupled Clip and Dynamic sAmpling Policy Optimization," arXiv v1 2025-03-18) states in its abstract that key technical details of top reasoning LLMs are concealed in both the o1 blog and the R1 report, then releases a fully open RL system built on the verl framework. Four techniques (Clip-Higher, dynamic sampling, token-level policy-gradient loss, overlong reward shaping) took Qwen2.5-32B-base to 50 points on AIME 2024 (arXiv 2503.14476), against R1-Zero-style 47 reported with the same base in the R1 report's comparison. Eight days later Sea AI Lab's "Understanding R1-Zero-Like Training" (2025-03-26) showed that Qwen2.5 base models already exhibit strong reasoning without a prompt template, and that GRPO has a length bias that inflates incorrect answers, proposing the corrected "Dr. GRPO" and 43.3% AIME 2024 with a 7B model (arXiv 2503.20783). This pair is the open community's answer to R1's missing RL code. It also introduced a caution (base-model priors matter) that Yue et al. and the "spurious rewards" paper extended. The lag is +57 days after R1.

Read it in the deep dive