Learning rewards from human feedback before deep RL, from 2008 to 2017

Researchers were learning rewards from people well before deep RL. TAMER ("Training an Agent Manually via Evaluative Reinforcement", W.

Date
August 2008
Who
UT Austin, INRIA, MIT Media Lab and others
Confidence
High (existence and citations); Medium (the reading that scale was the new part,
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): UT Austin, INRIA, MIT Media Lab and others · Confidence: High (existence and citations); Medium (the reading that scale was the new part, which is Inference) Researchers were learning rewards from people well before deep RL. TAMER ("Training an Agent Manually via Evaluative Reinforcement", W. Bradley Knox and Peter Stone, ICDL, August 2008) let a lay trainer give evaluative feedback, which the agent modeled and then exploited (paper page); the framework was restated in "Interactively Shaping Agents via Human Reinforcement: The TAMER Framework" (K-CAP, September 2009; paper page). Preference-based policy learning (Akrour et al. 2011/2012, Wilson et al. 2012) learned from pairwise comparisons, but with hand-coded features and, in Wilson's case, simulated humans; the Bradley-Terry model (1952), which Christiano et al. use to turn comparisons into a reward fit, is the statistical basis of today's reward-model loss (Christiano et al. 2017).

For language, Jaques et al. (2017) introduced the KL-control idea of keeping an RL-tuned sequence model close to its pretrained prior, which is where the later KL penalty comes from (Ziegler et al. 2019), and Kreutzer et al. (2018) and Böhm et al. (2019) used human feedback for translation and summarization (InstructGPT related work); Jaques et al. (2019) is the closest precedent to GPT-2 fine-tuning (B05-04a).

Inference: what was new in 2017 was scale (deep networks, no hand features, tiny clips). The concept already existed, and Christiano et al. themselves frame their contribution as scaling human feedback up to deep RL. Sources: TAMER, ICDL 2008 · TAMER framework, K-CAP 2009 · Christiano 2017 · Ziegler 2019 · Ouyang 2022

Read it in the deep dive