OpenAI fine-tunes GPT-2 from human preferences
An early documented use of explicit human comparisons plus PPO on a large pretrained language model (GPT-2 774M), with labeler shortcuts being exploited and the famous sign-flip bug.
- Date
- 18 September 2019
- Who
- OpenAI
- People
- Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, Geoffrey Irving
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Landmark · Significance: 4/5 · Org(s): OpenAI · People: Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, Geoffrey Irving · Confidence: High Primary sources: arXiv:1909.08593 (v1 2019-09-18) · OpenAI blog, 2019-09-19
One-liner. An early documented use of explicit human comparisons plus PPO on a large pretrained language model (GPT-2 774M), with labeler shortcuts being exploited and the famous sign-flip bug.
Why it happened. OpenAI had a pretrained model worth steering (GPT-2) and a safety team that wanted language, because "don't lie" cannot be expressed as an Atari score. The blog says language is also a necessary ingredient for amplification and debate (blog). The KL-control idea of Jaques et al. (2017) supplied the penalty that keeps the policy close to the pretrained model. Precedents and parallel work named in the paper's related-work section are Kreutzer et al. (2018, translation), Jaques et al. (2019, dialog, B05-04a), Yi et al. (2019, dialog) and, concurrently, Böhm et al. (2019, summarization). So "first" holds only for the combination of explicit pairwise comparisons, a large pretrained Transformer and PPO, and only as "first publicly documented" (Ziegler et al.).
The idea. Fine-tune the language model by RL against a reward model trained on human choices of the best of four samples.
How it works. The model is a 774M-parameter GPT-2 and the four tasks are positive-sentiment and physically-descriptive continuations of BookCorpus text, and summarization of Reddit TL;DR and CNN/Daily Mail. A reward model (initialized from the language model) is trained on human 4-way choices; the policy is optimized with PPO (B05-03) plus a KL penalty to the original model. Labels were collected online through Scale AI with a latency of about 30 minutes (paper, blog). This is the first appearance in the chapter of the commercial labeling vendor that later became central (B05-26, B05-42).
Results. Style tasks needed about 5,000 comparisons; labelers preferred the tuned model to zero-shot GPT-2 88% (sentiment) and 86% (descriptiveness) of the time. Summarization used 60,000 comparisons and the resulting policies were what the authors call "smart copiers" (their paper and blog use the phrase; the blog also says "smart copying engine"). They lifted whole sentences from the source, which labelers (and a simple lead-3 baseline) preferred to human reference summaries, though the authors judged the references better. The authors' own accuracy check on 30 articles per dataset found the tuned model accurate on 29/30 and 26/30, versus 6/30 for zero-shot (blog). Lab-measured.
How it spread. Directly into Stiennon et al. (B05-06) the next year; indirectly, the online-versus-batch data-collection lessons (quality control at low latency was hard; regressions were noticed only after runs) shaped later vendor practice.
Why it mattered. It showed the loop worked on language, and it surfaced the central failure modes early, namely labelers using cheap heuristics (copying) and optimization exploiting that.
Nuance, controversy and myths. The blog recounts a refactoring bug that flipped the sign of the reward and of the KL penalty. Because the instructions told labelers to rate sexually explicit text very low, the model optimized for exactly that content; the authors were asleep and found out only when the run finished. The failure produced fluent, "maximally bad" output, which is why it is a favorite interview anecdote (blog, "Bugs can optimize for bad behavior").
Interview kit.
- 30-second version: In 2019 OpenAI fine-tuned GPT-2 with human comparisons; it worked for sentiment with 5k labels and gave a copy-pasting summarizer with 60k, and a sign-flip bug accidentally produced the worst possible output.
- Likely follow-ups: Why did the summarizer copy? → Labelers were told to penalize inaccuracy, not copying, and copying is the easiest way to be accurate. Who labeled? → Scale AI contractors.
- Common mistake: Saying InstructGPT was the first RLHF on a language model (or, in the other direction, that Ziegler et al. were unambiguously first, since Jaques et al. and Kreutzer et al. came before).
- Connect it to: B05-02, B05-04a, B05-06, B05-18.
Sources. 1. Ziegler et al., arXiv:1909.08593 · 2. OpenAI, "Fine-tuning GPT-2 from human preferences", 2019-09-19