DeepMind and OpenAI propose reward modeling, debate and amplification for scalable oversight

This research programme held that human preference learning is only the first rung, and that when models become too capable for people to judge directly, models should help people judge.

Date
19 November 2018
Who
DeepMind, OpenAI
People
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, Shane Legg; Geoffrey Irving, Paul Christiano, Dario Amodei, Buck Shlegeris
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Landmark · Significance: 4/5 · Org(s): DeepMind, OpenAI · People: Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, Shane Legg; Geoffrey Irving, Paul Christiano, Dario Amodei, Buck Shlegeris · Confidence: High Primary sources: Leike et al., arXiv:1811.07871 (2018-11-19) · Irving et al., "AI safety via debate", arXiv:1805.00899 (2018-05-02) · Christiano et al., "Supervising strong learners by amplifying weak experts", arXiv:1810.08575 (2018-10-19)

One-liner. This research programme held that human preference learning is only the first rung, and that when models become too capable for people to judge directly, models should help people judge.

Why it happened. After the 2017 result (B05-02) the open question was whether it would work as tasks got harder, because preference learning works only while a human can tell which behavior is better. Leike's DeepMind paper frames this as the "agent alignment problem" and proposes reward modeling (learn the reward from user interaction, optimize it with RL) as a direction, and lists challenges such as reward hacking and the difficulty of evaluating complex outcomes (paper). Two OpenAI papers proposed siblings. Debate (Irving, Christiano and Amodei) has two models argue while a human judges, and iterated amplification (Christiano, Shlegeris and Amodei) decomposes a hard task into pieces that weaker helpers can supervise. Debate's abstract states the problem directly, saying preference judging "can fail if the task is too complicated for a human to directly judge" (paper).

The idea. Recursive reward modeling trains agent A1 with a reward model from human feedback, then uses A1 as an assistant to help the human evaluate a harder task for agent A2, and so on, so the evaluation capability bootstraps with the agent (paper §3.2). This technique is the recursive one with a human-feedback component, which ordinary RLHF is not (see the Terminology box).

Results. These were research directions with no experimental results in 2018. Empirical first steps came with book summarization (B05-09) and, at the frontier-lab level, the 2023 Superalignment announcement (B05-36).

How it spread. The people moved. Leike (DeepMind) later led alignment at OpenAI; Irving (OpenAI, then DeepMind) is the senior author on Sparrow (B05-16); Christiano left OpenAI for the Alignment Research Center, and Askell for Anthropic, by March 2022 (affiliations in the InstructGPT footnote). The lineage into scalable oversight and weak-to-strong research is in B22.

Why it mattered. It supplies the motive for RLHF at the safety teams, which wanted an oversight method that might scale. InstructGPT's own text calls RLHF a building block of these proposals (§5.1).

Nuance, controversy and myths. Debate and amplification were not shown to work at the time; Christiano describes RLHF as only a first step (Dwarkesh, 2023-10-31).

Interview kit.

  • 30-second version: RLHF only works while humans can judge outputs. This 2018 agenda proposed recursive reward modeling, debate and amplification so models can help humans judge harder tasks.
  • Likely follow-ups: Is that what "recursive learning with human feedback" means? → The nearest real thing is recursive reward modeling; ordinary RLHF is not recursive (Terminology). Did it work? → Partially, in toy and summarization settings; unresolved at scale (B22).
  • Common mistake: Treating RLHF and recursive self-improvement as the same thing; the latter is B24.
  • Connect it to: B05-02, B05-09, B05-36.

Sources. 1. Leike et al. 2018 · 2. Irving et al. 2018 · 3. Christiano et al. 2018 · 4. Ouyang et al. 2022

Read it in the deep dive