Scaling Laws for Reward Model Overoptimization, where pushing a proxy reward too far lowers the true one

Optimizing a reward model too hard makes the real objective worse. Gao, Schulman and Hilton quantified this by replacing human labelers with a fixed "gold" reward model (the 6B reward model from…

Date
19 October 2022
Who
OpenAI
People
Leo Gao, John Schulman, Jacob Hilton
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · People: Leo Gao, John Schulman, Jacob Hilton · Confidence: High Optimizing a reward model too hard makes the real objective worse. Gao, Schulman and Hilton quantified this by replacing human labelers with a fixed "gold" reward model (the 6B reward model from InstructGPT) that labels data to train smaller "proxy" reward models (3M to 3B parameters); they then optimize the policy against the proxy by RL or best-of-n sampling and track the gold score as a function of KL distance from the initial policy. The gold score rises then falls, follows different functional forms for best-of-n and RL (the RL form is d(α − β log d), where d is the square root of the KL), and the fit coefficients scale smoothly with reward-model size (arXiv:2210.10760, v1 2022-10-19). OpenAI's ChatGPT post cites this paper for the verbosity and phrase-overuse problems it saw (ChatGPT post). It also explains why the KL penalty exists and why Karpathy later called the reward model "just a vibe check" you cannot optimize too long, adding that no convincing open-domain RL on LLMs had been shown (Karpathy's 2024-08-08 post, quoted by Willison). Caveat: the "gold" is itself a model, so the paper measures how a proxy diverges from another model and does not measure divergence from human judgments. Sources: Gao et al. · ChatGPT post · Karpathy via Willison

Read it in the deep dive