Jaques et al. at MIT use a pretrained prior, KL-control and human-preference rewards in dialog
Jaques et al. at MIT published the closest precedent to GPT-2 fine-tuning, 80 days before it.
- Date
- 30 June 2019
- Who
- MIT (Media Arts and Science)
- People
- Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, Rosalind Picard
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): MIT (Media Arts and Science) · People: Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, Rosalind Picard · Confidence: High Jaques et al. at MIT published the closest precedent to GPT-2 fine-tuning, 80 days before it. Their paper "Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog" used a pretrained model as a prior and KL-control to penalize divergence from that prior during RL, with rewards extracted from human interaction data (implicit signals such as sentiment and engagement, with no explicit pairwise comparisons), on open-domain dialog generation with a 20,000-dimensional action space, and tested the result by deploying the bots live (arXiv:1907.00456, v1 2019-06-30). It is the reason "first RLHF on a language model" is contested. Ziegler et al. cite it as human-evaluation-as-reward work and credit Jaques et al. (2017; 2019) for the KL penalty (Ziegler et al.). What was new in Ziegler et al. was the combination of explicit four-way human comparisons, a large pretrained Transformer and PPO (B05-05). Sources: Jaques et al. · Ziegler et al.