Scaling Laws for Reward Model Overoptimization, where pushing a proxy reward too far lowers the true one
Optimizing a reward model too hard makes the real objective worse. Gao, Schulman and Hilton quantified this by replacing human labelers with a fixed "gold" reward model (the 6B reward model from…
- Date
- 19 October 2022
- Who
- OpenAI
- People
- Leo Gao, John Schulman, Jacob Hilton
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · People: Leo Gao, John Schulman, Jacob Hilton · Confidence: High Optimizing a reward model too hard makes the real objective worse. Gao, Schulman and Hilton quantified this by replacing human labelers with a fixed "gold" reward model (the 6B reward model from InstructGPT) that labels data to train smaller "proxy" reward models (3M to 3B parameters); they then optimize the policy against the proxy by RL or best-of-n sampling and track the gold score as a function of KL distance from the initial policy. The gold score rises then falls, follows different functional forms for best-of-n and RL (the RL form is d(α − β log d), where d is the square root of the KL), and the fit coefficients scale smoothly with reward-model size (arXiv:2210.10760, v1 2022-10-19). OpenAI's ChatGPT post cites this paper for the verbosity and phrase-overuse problems it saw (ChatGPT post). It also explains why the KL penalty exists and why Karpathy later called the reward model "just a vibe check" you cannot optimize too long, adding that no convincing open-domain RL on LLMs had been shown (Karpathy's 2024-08-08 post, quoted by Willison). Caveat: the "gold" is itself a model, so the paper measures how a proxy diverges from another model and does not measure divergence from human judgments. Sources: Gao et al. · ChatGPT post · Karpathy via Willison