Jaques et al. at MIT use a pretrained prior, KL-control and human-preference rewards in dialog

Jaques et al. at MIT published the closest precedent to GPT-2 fine-tuning, 80 days before it.

Date
30 June 2019
Who
MIT (Media Arts and Science)
People
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, Rosalind Picard
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): MIT (Media Arts and Science) · People: Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, Rosalind Picard · Confidence: High Jaques et al. at MIT published the closest precedent to GPT-2 fine-tuning, 80 days before it. Their paper "Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog" used a pretrained model as a prior and KL-control to penalize divergence from that prior during RL, with rewards extracted from human interaction data (implicit signals such as sentiment and engagement, with no explicit pairwise comparisons), on open-domain dialog generation with a 20,000-dimensional action space, and tested the result by deploying the bots live (arXiv:1907.00456, v1 2019-06-30). It is the reason "first RLHF on a language model" is contested. Ziegler et al. cite it as human-evaluation-as-reward work and credit Jaques et al. (2017; 2019) for the KL penalty (Ziegler et al.). What was new in Ziegler et al. was the combination of explicit four-way human comparisons, a large pretrained Transformer and PPO (B05-05). Sources: Jaques et al. · Ziegler et al.

Read it in the deep dive