Ai2 introduced "RLVR", reinforcement learning with verifiable rewards, in Tülu 3

Ai2's open post-training report (arXiv v1 2024-11-22) introduces "Reinforcement Learning with Verifiable Rewards", which keeps the RLHF objective and replaces the learned reward model with a…

Date
22 November 2024
Who
Allen Institute for AI (Ai2), University of Washington
People
Nathan Lambert, Hannaneh Hajishirzi, Valentina Pyatkin et al.
Confidence
High
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): Allen Institute for AI (Ai2), University of Washington · People: Nathan Lambert, Hannaneh Hajishirzi, Valentina Pyatkin et al. · Confidence: High Ai2's open post-training report (arXiv v1 2024-11-22) introduces "Reinforcement Learning with Verifiable Rewards", which keeps the RLHF objective and replaces the learned reward model with a deterministic checker that pays a fixed reward when a completion is verifiably correct. It trains with PPO on grade-school math, MATH and precise-instruction-following constraints, on Llama 3.1 8B and 70B bases in v1; the 405B results were added in v5 on 2025-04-14 (Tülu 3, §6; version history on the arXiv page). The paper presents it as a simplification of earlier bootstrapping work (STaR) and makes no claim about long-chain-of-thought emergence. It supplied the name the field adopted ("RLVR") for exactly what R1-Zero scaled two months later. Lineage and the STaR/process-supervision precursors are in B07, and the preference-tuning side is in B06.

Read it in the deep dive