OpenAI trains a 1.3B summarizer from human feedback that beats a 10x larger supervised model

A 1.3B model trained with human feedback beat a supervised model ten times its size and the human reference summaries on Reddit TL;DR, with the full three-step pipeline that InstructGPT later reused.

Date
2 September 2020
Who
OpenAI
People
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, Paul Christiano
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Landmark · Significance: 4/5 · Org(s): OpenAI · People: Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, Paul Christiano · Confidence: High Primary sources: arXiv:2009.01325 (v1 2020-09-02) · NeurIPS 2020

One-liner. A 1.3B model trained with human feedback beat a supervised model ten times its size and the human reference summaries on Reddit TL;DR, with the full three-step pipeline that InstructGPT later reused.

Why it happened. Ziegler et al. (B05-05) had left two problems, noisy labels and weak base models. This paper made two changes to improve data quality. It moved from Ziegler's online loop to batched collection, alternating between sending large batches of comparisons to labelers and retraining on the cumulative data; and it kept a hands-on relationship with labelers (detailed onboarding, a shared chat room for questions, regular feedback, continuous monitoring of labeler-researcher agreement). It also used much larger models (1.3B and 6.7B, GPT-3-style architectures).

The idea. (1) Collect samples from existing policies and send comparisons to humans; (2) train a reward model on them; (3) optimize a policy with PPO against that reward model, with a KL penalty to the supervised model (paper). Labelers were recruited from Upwork and two labeling services, Scale and Lionbridge.

How it works. 64,832 human summary comparisons (released as a dataset). Researcher-labeler agreement was about 77%, versus 73% between researchers; the reward model grew more accurate with both data and size (doubling data gave about +1.1 points of validation accuracy; doubling model size about +1.8). RL fine-tuning of the 6.7B model took about 320 GPU-days.

Results. The 1.3B human-feedback model was preferred over reference summaries 61% of the time versus 43% for a supervised model ten times larger; the 6.7B model beat both and, after controlling for length (the models wrote longer summaries), was still preferred about 65% of the time. Transfer to CNN/Daily Mail news without news-specific tuning gave summaries nearly as good as the references (company-measured, with human evaluation).

How it spread. Within OpenAI, Ouyang, Wu, Lowe and Christiano carried over to the book-summarization, WebGPT and InstructGPT papers (B05-09, B05-11, B05-13). Anthropic's 2021 paper used the summarization comparisons as a transfer benchmark (B05-10) and Llama 2's reward models were trained partly on the released OpenAI Summarize data (B05-37).

Why it mattered. It showed that RLHF beats imitation, because a smaller RL-tuned model outperformed a larger model trained on human-written summaries. It also contains an early careful measurement of over-optimization. Past a point, pushing the reward model higher made summaries worse by labeler preference, and eventually the reward model was anti-correlated with true preference (paper, quantified later in B05-18).

Nuance, controversy and myths. The authors flag their cost (thousands of labeler hours) and that labelers preferred lead-3 to references on CNN/DM, i.e. the "ground truth" is itself a human-judgment artifact.

Interview kit.

  • 30-second version: Before InstructGPT, OpenAI proved the recipe on summarization with SFT, then a reward model from 64k comparisons, then PPO. A 1.3B RLHF model beat a 13B supervised one.
  • Likely follow-ups: Why is that surprising? → Optimizing a learned preference signal gave more than 10x parameters of imitation. What goes wrong when you optimize too hard? → Reward-model over-optimization (B05-18).
  • Common mistake: Calling it "ChatGPT's paper"; it is a summarization system.
  • Connect it to: B05-05, B05-13.

Sources. 1. Stiennon et al., arXiv:2009.01325

Read it in the deep dive