OpenAI's "Let's Verify Step by Step", which introduces process supervision and the PRM800K dataset
OpenAI trained reward models on 800,000 human step-by-step correctness labels for math solutions and showed that rewarding each step beats rewarding only the final answer.
- Date
- 31 May 2023
- Who
- OpenAI
- People
- Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Landmark · Significance: 4/5 · Org(s): OpenAI · People: Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe · Confidence: High Primary sources: arXiv:2305.20050 (v1 2023-05-31) · PRM800K dataset
One-liner. OpenAI trained reward models on 800,000 human step-by-step correctness labels for math solutions and showed that rewarding each step beats rewarding only the final answer. It applies the human-feedback loop from RLHF to reward models for reasoning.
Why it happened. RLHF reward models score a whole response by human preference (B05-13). For multi-step reasoning there are two ways to give feedback, outcome supervision (judge the final answer) and process supervision (judge each step). The paper's framing is cost. Human feedback is expensive, so the two should be compared carefully; for the MATH dataset outcome supervision can be automated because answers are checkable, while process supervision needs human labelers. DeepMind's Uesato et al. (2022) had defined the two methods and found similar final performance on grade-school math; the authors repeat the comparison on the harder MATH dataset at larger scale. Leike and Schulman, who appear throughout this chapter, are co-authors (paper).
The idea. Train a reward model that scores each step of a solution, then use it to pick among many sampled solutions.
How it works. The large-scale models are fine-tuned from the base GPT-4 model, which the paper says was pretrained only to predict the next token and had no RLHF, plus an extra pass on roughly 1.5B math-relevant tokens; the generator is taught to write newline-delimited steps. Human data-labelers saw step-by-step solutions to MATH problems and marked each step positive, negative or neutral. After filtering, the dataset, released as PRM800K, has about 800,000 step-level labels over 75,000 solutions, collected in two phases (about 5% in a first, cumbersome phase); the second phase, most of the data, used active learning, in which the current reward model ranked sampled solutions and the highest-scoring wrong-answer ones were shown to labelers over 10 generations. The authors report that this significantly improves efficacy. Small-scale ablations used a large reward model as a stand-in for human supervision.
Results. On a representative subset of the MATH test set, choosing the best of 1,860 samples with the process-supervised reward model solved 78.2% of problems, against 72.4% for an outcome-supervised reward model and 69.6% for majority voting (Figure 3). Lab-measured, on one benchmark.
How it spread.
| Lab/project | Response | Relationship | Date | Lag vs. original |
|---|---|---|---|---|
| DeepSeek-R1 report | Cites the paper, then reports that process reward models have three practical limits (hard to define a fine-grained step, hard to judge intermediate correctness at scale, reward hacking once a model-based reward is used) and that their advantage was limited against the added compute in large-scale RL; its recipe leans on rule-based rewards (arXiv:2501.12948, §4.2) | critique and alternative | 2025-01-22 | 20 months |
Most of the follow-on story belongs to the reasoning chapters (B07, B08).
Why it mattered. It moved the human-feedback loop from "which answer do you prefer" to "is this step correct", created a reusable labeled dataset, and made the labor model of RLHF (many people judging) concrete for reasoning. Inference: the later shift to rule-based, verifiable rewards in reasoning RL can be read partly as a way to avoid this labor cost (B08).
Nuance, controversy and myths. (1) This is human feedback but not preference RLHF. The labels are about correctness, the reward model was evaluated by reranking samples (best-of-N), and the paper says using it for RL is intentionally not its focus. (2) The result is on MATH; whether step labels generalize to open-ended tasks is exactly what DeepSeek later questioned. (3) OpenAI has not, to my knowledge, said how o1 was trained in this detail; do not state that o1 used PRM800K.
Interview kit.
- 30-second version: OpenAI showed that a reward model trained on human labels for every step of a math solution beats one trained only on final answers (78.2% versus 72.4% best-of-1,860 on MATH) and released 800,000 step labels.
- Likely follow-ups: Is that RLHF? → It is human-feedback reward modeling with correctness labels instead of preference labels, used for reranking. Why did later reasoning RL mostly drop it? → Step labels are costly and hard to define for general tasks; verifiable outcome rewards are cheaper (B08).
- Common mistake: Equating process supervision with chain-of-thought prompting; it is a training-signal choice.
- Connect it to: B05-13, B05-18, B05-40c, B07, B08.
Sources. 1. Lightman et al., arXiv:2305.20050 · 2. DeepSeek-R1, arXiv:2501.12948