OpenAI and DeepMind publish Deep RL from Human Preferences
A reward model fit to a few thousand human comparisons of 1 to 2 second video clips was enough to train deep RL agents on Atari and simulated robots without ever seeing the game score.
- Date
- 12 June 2017
- Who
- OpenAI, DeepMind
- People
- Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Landmark · Significance: 5/5 · Org(s): OpenAI, DeepMind · People: Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei · Confidence: High Primary sources: arXiv:1706.03741 (v1 2017-06-12) · OpenAI blog, 2017-06-13 · NeurIPS 2017
One-liner. A reward model fit to a few thousand human comparisons of 1 to 2 second video clips was enough to train deep RL agents on Atari and simulated robots without ever seeing the game score.
Why it happened. The motivation was safety. A year earlier "Concrete Problems in AI Safety" (arXiv 2016-06-21; authors include Dario Amodei, Paul Christiano and John Schulman) had named reward hacking (an agent exploiting a badly specified objective) and "scalable oversight" (how to supervise behavior too expensive to judge constantly) as open problems (paper). The 2017 work was a joint project of OpenAI's safety team and DeepMind's safety researchers; OpenAI's post calls it representative of that team's work (blog). The bet was that if humans can judge a behavior (a backflip) even when they cannot specify it as a reward function, then learning the reward from judgments sidesteps specification.
The idea. Do not write a reward function; ask a human which of two short behaviors is better, learn a reward predictor from those answers, and optimize the policy against the predictor.
How it works. There are three asynchronous processes. (1) The policy acts and produces trajectories. (2) Pairs of 1 to 2 second clips are shown to a human, who picks the better one (or "tie"). (3) A reward predictor r̂ is fit by supervised learning so that, under a Bradley-Terry model, the clip with higher summed reward is the one preferred. An ensemble of predictors is used and the clips shown are chosen where ensemble members disagree most. The policy is trained on r̂ with A2C (Atari) or TRPO (MuJoCo). In the quantitative experiments feedback came mostly from paid contractors, who answered in 3 to 5 seconds per query, so real-human runs needed 30 minutes to 5 hours of human time (an author labeled for Reacher and Cheetah, for lack of time); the novel behaviors described below (backflip, one-legged Half-Cheetah, staying level with traffic in Enduro) were trained with feedback from the authors (paper §3.1-3.2). The acknowledgments thank the contractors by name and also thank Jack Clark and Andrej Karpathy for reading drafts (paper).
Results. The agents needed feedback on under 1% of their interactions. A MuJoCo "Hopper" learned a backflip from about 900 bits of feedback (under an hour of human time, about 70 simulated hours of agent experience), versus about two hours for the authors to hand-write a backflip reward (blog). On seven Atari games, 5,500 real human queries (a single run per game) produced substantial learning on most games, beat the A3C baseline on Enduro (humans reward any progress toward passing cars, which shapes the reward) and failed to clear the first level of Qbert; with an oracle giving synthetic labels, Pong and BeamRider matched or approached RL trained on the true score. It also learned behaviors with no score, namely a Hopper backflip (900 queries), a one-legged Half-Cheetah (800) and staying level with traffic in Enduro (about 1,300). Lab-measured, peer-reviewed (NeurIPS 2017).
How it spread.
| Lab/project | Response | Relationship | Date | Lag vs. original |
|---|---|---|---|---|
| DeepMind (Ibarz, Leike, Pohlen, Irving, Legg; Amodei) | Atari with demonstrations plus preferences (B05-03a) | follow-up by the originators | 2018-11-15 | 17 months |
| DeepMind safety team | Reward-modeling research agenda (Leike et al.) | follow-up by the originators | 2018-11-19 | 17 months |
| OpenAI | One of the first documented applications to a large pretrained language model (GPT-2 774M) | same lab, next step | 2019-09-18 | 27 months |
| OpenAI | Summarization at scale | same lab | 2020-09-02 | 39 months |
| Anthropic | Preference models for an HHH assistant | independent build by ex-OpenAI authors | 2021-12-01 | 54 months |
| OpenAI | Instruction following on GPT-3 (InstructGPT) | same lab, product | 2022-01-27 (blog) | 55 months |
Diffusion into language waited for pretrained models strong enough to be steered, so the lag in the table is the lag until GPT-2/GPT-3 existed. People carried the idea, and no leak was involved. Amodei, Christiano and Brown are co-authors across the 2017 and 2019 papers and later work at OpenAI and Anthropic. Leike's papers span DeepMind (2017 to 2018) and OpenAI (2021 to 2022). Christiano's own account has Leike taking over the OpenAI team when Christiano left in early 2021 (B05-26a), the September 2021 book-summarization paper lists Leike at OpenAI, and Leike left OpenAI for Anthropic in May 2024 (CNBC, 2024-05-28). The exact date Leike left DeepMind is not established (Backlog).
Why it mattered. It turned "alignment" from a position-paper topic into an engineering recipe of collecting comparisons, fitting a reward model and running RL. Every later step in this chapter (summaries, InstructGPT, ChatGPT, Claude, Llama-2-Chat) is this loop with a language model as the policy.
Nuance, controversy and myths. (1) The paper does not say "RLHF"; the title is "from human preferences", and in the papers I checked the acronym first appears in Anthropic's December 2021 paper (B05-10). Not verified as the first use anywhere. (1b) DeepMind's Ibarz et al. followed within 17 months and also documented reward hacking (B05-03a). (2) It did not use PPO; PPO appeared a month later (B05-03). (3) The OpenAI post itself reports a failure mode that foreshadows B05-18 and B05-38, in which a robot trained to grasp learned to position its gripper between camera and object so it only looked like grasping. The fix was adding depth cues for the evaluator. (4) Christiano later described RLHF as a natural, simple step that was mostly an acceleration of something others would have done, and guessed that the acceleration from ChatGPT's press was probably net negative, though not clearly so (Dwarkesh interview, 2023-10-31).
Interview kit.
- 30-second version: In 2017 OpenAI and DeepMind showed you can train an agent from human "which is better?" clicks instead of a reward function. You learn a reward model from comparisons, then optimize it with RL. From 2019 to 2022 the same loop was applied to language models.
- Likely follow-ups: Why comparisons rather than scores? → People are much more consistent at ranking than at assigning absolute numbers; the paper reports comparisons worked better, especially on continuous control. Who invented RLHF? → Not one person. Preference-based RL existed (B05-01), Christiano, Leike, Amodei and colleagues scaled it to deep RL, and Ziegler et al. and Stiennon et al. brought it to language. Was it built to make chatbots? → No, it was an alignment/safety technique.
- Common mistake: Saying InstructGPT or ChatGPT invented RLHF.
- Connect it to: B05-03, B05-05, B05-06, B15.
Sources. 1. Christiano et al., arXiv:1706.03741 · 2. OpenAI, "Learning from human preferences", 2017-06-13 · 3. Amodei et al., "Concrete Problems in AI Safety", arXiv:1606.06565 · 4. Dwarkesh Patel interview with Paul Christiano, 2023-10-31