OpenAI releases InstructGPT, GPT-3 tuned with human demonstrations, rankings and PPO

Human demonstrations, human rankings and PPO made a 1.3B-parameter model preferable to the 175B base GPT-3 for real user prompts; this is the recipe behind ChatGPT.

Date
27 January 2022
Who
OpenAI (Alignment team)
People
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin (lead authors); Ryan Lowe and Jan Leike (team leads); John Schulman; Paul Christiano; Amanda Askell
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Landmark · Significance: 5/5 · Org(s): OpenAI (Alignment team) · People: Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin (lead authors); Ryan Lowe and Jan Leike (team leads); John Schulman; Paul Christiano; Amanda Askell · Confidence: High Primary sources: OpenAI blog, 2022-01-27 · arXiv:2203.02155 (v1 2022-03-04; NeurIPS 2022)

One-liner. Human demonstrations, human rankings and PPO made a 1.3B-parameter model preferable to the 175B base GPT-3 for real user prompts; this is the recipe behind ChatGPT.

Why it happened. OpenAI had a customer-facing GPT-3 API (private beta from 2020-06-11, whose launch post already mentioned learning from human feedback provided by users or labelers, OpenAI API post) and a measurable gap. People wanted models to follow instructions, but GPT-3 was trained to predict the next token on internet text, so it made up facts, was toxic, or ignored the request. The alignment team (the paper says it was a joint project of the OpenAI Alignment team, with Lowe and Leike as team leads) had the RLHF skills from the 2017 to 2020 work (B05-02, B05-05, B05-06) and a stream of real prompts, since the training prompts came from customers' use of earlier InstructGPT models in the API Playground, with a notice that data could be used for training. The blog says InstructGPT models had been in beta on the API for more than a year, and the earliest Playground version, trained on demonstrations only, had been deployed in January 2021 (blog, footnote; paper §3.2). The motive was both product (a better API) and safety (the first time OpenAI's alignment research "has been applied to our product", in the blog's words). OpenAI's 2022-08-24 statement of its alignment approach spells out the logic. The natural-language API is a useful environment for alignment research because it supplies a rich real-world feedback loop on tasks customers will pay for; RLHF was then its main technique for deployed models; and it saw RLHF as a core building block for scalable-oversight proposals, though it did not expect RLHF alone to be sufficient for AGI (OpenAI, 2022-08-24).

The idea. Align the model to what users intend, defined as helpful, honest, harmless (the HHH vocabulary of B05-10), by learning from labeler-written demonstrations and labeler rankings of outputs.

How it works. There are three steps. (1) SFT fine-tunes GPT-3 on labeler demonstrations (about 13k training prompts, 16 epochs even though validation loss rose after one epoch because human-preference scores kept improving). (2) Reward model. A 6B model (starting from the SFT model with the unembedding layer removed) was trained on rankings, with labelers ranking K=4 to 9 outputs per prompt (about 33k training prompts), and all K-choose-2 pairs of a prompt were trained as one batch element to avoid overfitting. A 175B reward model was unstable. (3) PPO optimizes the SFT model against the reward model (B05-03) on about 31k API-only prompts, with a per-token KL penalty. PPO-ptx (the final "InstructGPT") mixes in log-likelihood updates on the pretraining distribution to cut "alignment tax" regressions on SQuAD, DROP, HellaSwag and WMT 2015 fr-en (paper §1, §3.5). The blog's mental model is that RLHF brings out abilities already in GPT-3, using under 2% of pretraining compute and data, and the August 2022 post adds that InstructGPT's fine-tuning cost under 2% of GPT-3's pretraining compute and about 20,000 hours of human feedback (OpenAI, 2022-08-24). In the paper's numbers, 175B SFT cost 4.9 and 175B PPO-ptx 60 petaflop/s-days versus 3,640 for GPT-3.

Labor. About 40 contractors hired through Upwork and Scale AI after a screening test (agreement with researchers on sensitive-speech flagging and rankings, a demonstration-writing score, soft cutoffs of 75% agreement and 6/7); labelers agreed with each other 72.6% of the time (77.3% for held-out labelers); most comparisons were labeled by a single contractor because of cost; labelers lived mostly in the US or Southeast Asia; they prioritized helpfulness during training but truthfulness and harmlessness in final evaluations; demographics come from a voluntary survey with 19 respondents, who said they enjoyed the task and felt fairly paid (self-reported); the acknowledgments thank about 40 labelers by name, as the 2017 paper did for its contractors (paper §3.4, App. B).

Results. Labelers preferred the 175B InstructGPT to 175B GPT-3 85 ± 3% of the time and to few-shot-prompted GPT-3 71 ± 4%; the 1.3B PPO-ptx model was preferred to 175B GPT-3. On TruthfulQA the model was about twice as often truthful and informative. Hallucination on closed-domain tasks was 21% versus 41%, there were about 25% fewer toxic outputs when prompted to be respectful, and there was no improvement on bias benchmarks (Winogender, CrowS-Pairs). It also followed instructions in other languages and about code despite little such training data. All measured by OpenAI on its own labelers and customer-style prompts.

How it spread.

Lab/projectResponseRelationshipDateLag vs. blog
OpenAIDeployed as default API models (blog)own product2022-01-270
DeepMindGopherCite, 280B Gopher trained with RL from human preferences (B05-13a)parallel, independent line (DeepMind's own RLHF work since 2018)2022-03-212 months
AnthropicHH-RLHF paper with PMs and RL at 52B (B05-14)independent build by ex-OpenAI authors (HHH paper predates the InstructGPT blog)2022-04-122.5 months
DeepMindSparrow, RLHF with rules (B05-16)parallel, independent line2022-09-228 months
Open-sourcetrlX (repo 2022-10-03) and RL4LMs (paper 2022-10-03) open RLHF code for larger models; the TRL library (PPO for Hugging Face models) dates from 2020 (B05-17)derived2022-10-038 months
OpenAIChatGPT, "same methods as InstructGPT" (B05-20)same lab, sibling product2022-11-3010 months
MetaLlama 2-Chat with 1.4M comparisons (B05-37)derived (documents and extends the recipe)2023-07-1818 months
GoogleGemini describes SFT then RLHF (B05-40)own pipeline, documented late2023-12-0622 months

Diffusion was fast inside labs that had alignment researchers and slow elsewhere. It was gated by (a) human-data pipelines and money for labelers, (b) a PPO implementation that needs several model copies in memory, and (c) base-model quality. OpenAI published the method, so no leak or distillation was needed; open-source RLHF reproductions lagged because of (a) and (b) (B05-17).

Why it mattered. It showed that alignment is capability for users, because a method designed as a safety technique made models more useful than a 100x scale-up. The authors say as much and call RLHF a more cost-effective way to improve user-perceived quality than larger models (§5.1).

Nuance, controversy and myths. (1) 1.3B vs 175B is vs base GPT-3, not vs the 175B InstructGPT. (2) The paper's own limitations are that it aligns to the preferences of about 40 contractors, the researchers' instructions and API customers, not to a broader notion of human values, and that models follow harmful instructions and, when asked to be maximally biased, produce more toxic output than GPT-3 (§5.2-5.3). (3) The API models called InstructGPT (text-davinci-001/002) were trained on the same human data but with a similar-but-different method that, as it turned out, was not PPO RLHF (B05-19). (4) It is not "ChatGPT's paper", because ChatGPT adds dialogue data and a chat interface (B05-20).

Interview kit.

  • 30-second version: OpenAI's Alignment team took GPT-3, trained it on human-written demonstrations, trained a reward model on human rankings, and optimized with PPO. Labelers preferred the 1.3B result to the 175B original.
  • Likely follow-ups: Why did a smaller model win? → The base model already had the knowledge; RLHF changed what it does with it, matching what users want. Who did the labeling? → About 40 contractors via Upwork and Scale, mostly US and Southeast Asia. What's the catch? → Alignment to a small group's preferences, an alignment tax fixed with PPO-ptx, and still toxic/hallucinating.
  • Common mistake: Saying InstructGPT was trained on "human feedback" without the three-step structure, or that it is ChatGPT.
  • Connect it to: B05-06, B05-14, B05-19, B05-20, B03.

Sources. 1. OpenAI, "Aligning language models to follow instructions", 2022-01-27 · 2. Ouyang et al., arXiv:2203.02155 · 3. OpenAI, "Introducing ChatGPT"

Read it in the deep dive