Ai2 introduced "RLVR", reinforcement learning with verifiable rewards, in Tülu 3
Ai2's open post-training report (arXiv v1 2024-11-22) introduces "Reinforcement Learning with Verifiable Rewards", which keeps the RLHF objective and replaces the learned reward model with a…
- Date
- 22 November 2024
- Who
- Allen Institute for AI (Ai2), University of Washington
- People
- Nathan Lambert, Hannaneh Hajishirzi, Valentina Pyatkin et al.
- Confidence
- High
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): Allen Institute for AI (Ai2), University of Washington · People: Nathan Lambert, Hannaneh Hajishirzi, Valentina Pyatkin et al. · Confidence: High Ai2's open post-training report (arXiv v1 2024-11-22) introduces "Reinforcement Learning with Verifiable Rewards", which keeps the RLHF objective and replaces the learned reward model with a deterministic checker that pays a fixed reward when a completion is verifiably correct. It trains with PPO on grade-school math, MATH and precise-instruction-following constraints, on Llama 3.1 8B and 70B bases in v1; the 405B results were added in v5 on 2025-04-14 (Tülu 3, §6; version history on the arXiv page). The paper presents it as a simplification of earlier bootstrapping work (STaR) and makes no claim about long-chain-of-thought emergence. It supplied the name the field adopted ("RLVR") for exactly what R1-Zero scaled two months later. Lineage and the STaR/process-supervision precursors are in B07, and the preference-tuning side is in B06.