Open Problems and Fundamental Limitations of RLHF
A 32-author survey led by Stephen Casper and Xander Davies sorts RLHF's problems into tractable ones (fixable within RLHF) and fundamental ones (requiring alternatives).
- Date
- 27 July 2023
- Who
- MIT CSAIL, Harvard, Berkeley, Stanford, Cornell Tech, Apollo Research, ETH Zurich and others
- People
- Stephen Casper and Xander Davies (equal-contribution leads); Dylan Hadfield-Menell (senior) and about 30 co-authors including Anca Dragan, David Krueger, Dorsa Sadigh
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Landmark · Significance: 4/5 · Org(s): MIT CSAIL, Harvard, Berkeley, Stanford, Cornell Tech, Apollo Research, ETH Zurich and others · People: Stephen Casper and Xander Davies (equal-contribution leads); Dylan Hadfield-Menell (senior) and about 30 co-authors including Anca Dragan, David Krueger, Dorsa Sadigh · Confidence: High Primary sources: arXiv:2307.15217 (v1 2023-07-27; v2 2023-09-11)
One-liner. A 32-author survey led by Stephen Casper and Xander Davies sorts RLHF's problems into tractable ones (fixable within RLHF) and fundamental ones (requiring alternatives). It argues that RLHF needs auditing and multi-layered safety measures around it.
Why it happened. By mid-2023 RLHF was the central post-training method at every frontier lab (B05-37) yet, in the authors' words, there was little public work systematizing its flaws. The authors came largely from academic alignment and RL groups outside the labs.
The idea. A taxonomy across three stages (the human feedback, the reward model and the policy), plus joint training of the reward model and policy (Fig. 1).
How it works. The paper lists these fundamental problems. Humans cannot evaluate performance on difficult tasks, and can be misled so their evaluations are gamed. There is an inherent cost-quality trade-off in collecting feedback. One person's values are hard to represent with a reward function, and a single reward function cannot represent a diverse society. Reward models can misgeneralize, and optimizing an imperfect proxy leads to reward hacking. Policies can fail in deployment even if training rewards were fine, and optimal RL agents tend to seek power. Tractable problems include annotator biases and data poisoning, evaluating reward models, adversarial exploitability and mode collapse from RL fine-tuning. The paper also proposes disclosure and auditing standards for RLHF systems.
Results. The paper reports no experiments. It is an organizing document with a long reading list, and its main message is that RLHF is a useful but incomplete tool that should be combined with other safeguards.
How it spread. I did not measure its citations. The failure modes it names show up in empirical work on sycophancy (B05-39), over-optimization (B05-18) and the 2025 GPT-4o episode (B05-41).
Why it mattered. It converted a pile of known issues into an agenda and gave interviewers and policymakers a shared vocabulary for "what is wrong with RLHF".
Nuance, controversy and myths. The tractable/fundamental labels are judgments, and later work has moved some "fundamental" items (e.g., human evaluation limits) into active research via AI-assisted oversight (B22). The paper does not claim RLHF is useless.
Interview kit.
- 30-second version: Casper et al. (2023) catalog what is wrong with RLHF, including that humans can't judge hard tasks and can be fooled, that one reward can't represent everyone, reward hacking and mode collapse. They propose auditing and layered defenses.
- Likely follow-ups: "Which problems are fundamental?" → Evaluation of hard tasks, misleading evaluators, diverse values, reward misspecification and hacking. "Is it anti-RLHF?" → No. The paper argues against relying on RLHF alone.
- Common mistake: Describing it as an empirical result when it is a survey.
- Connect it to: B05-18, B05-25, B05-39, B22.
Sources. 1. Casper et al., arXiv:2307.15217