Four Habits, Open-Reasoner-Zero and OpenThoughts fill in what R1 left out
Three open papers filled in what R1 left out. (1) Gandhi et al., "Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs" (arXiv v1 2025-03-03, +42 days…
- Date
- 3 March 2025
- Who
- academic and open-source groups
- Confidence
- High (all three papers opened on arXiv)
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): academic and open-source groups · Confidence: High (all three papers opened on arXiv) Three open papers filled in what R1 left out. (1) Gandhi et al., "Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs" (arXiv v1 2025-03-03, +42 days after R1), asked why identical RL on the Countdown game lifts Qwen-2.5-3B far above Llama-3.2-3B. The answer was that Qwen already shows four behaviours (verification, backtracking, subgoal setting, backward chaining), that priming Llama with examples of them lets RL match Qwen, and that the presence of the behaviours mattered more than whether the example answers were right (arXiv 2503.01307). (2) Open-Reasoner-Zero (arXiv v1 2025-03-31, +70) reports that vanilla PPO with GAE (λ = 1, γ = 1) and simple rule-based rewards, with no KL regularisation, reproduces R1-Zero's growth in score and response length on the Qwen2.5-32B base, beating R1-Zero-Qwen-32B on AIME 2024, MATH-500 and GPQA Diamond in a tenth of the training steps, with code, data and weights released (arXiv 2503.24290). (3) OpenThoughts (arXiv v1 2025-06-04, +135) is an open data recipe. Over 1,000 controlled experiments on the data pipeline led to a 1.2M-example dataset with QwQ-32B as teacher and a 7B model scoring 53% on AIME 2025, 51% on LiveCodeBench and 54% on GPQA Diamond, all public (arXiv 2506.04178). They matter for different reasons. Gandhi et al. give a mechanism for the Qwen-only caveat raised by Dr. GRPO and Yue et al.. Open-Reasoner-Zero shows R1-Zero's scaling does not need GRPO's critic-free design. OpenThoughts shows the distillation route has an open data recipe, a counterpart to Sky-T1 and s1. Sources: arXiv 2503.01307 · arXiv 2503.24290 · arXiv 2506.04178