Google's SLiC-HF, a simpler alternative to PPO published twelve days before DPO
Sequence Likelihood Calibration with Human Feedback (SLiC-HF) adapted Google's earlier SLiC method to learn from human preference pairs, including preference data collected for a different model (an…
- Date
- 17 May 2023
- Who
- Google (the paper lists Google DeepMind and Google Research)
- People
- Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, Peter J. Liu
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): Google (the paper lists Google DeepMind and Google Research) · People: Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, Peter J. Liu · Confidence: High Sequence Likelihood Calibration with Human Feedback (SLiC-HF) adapted Google's earlier SLiC method to learn from human preference pairs, including preference data collected for a different model (an off-policy, offline setting), and on TL;DR summarization the authors report that it significantly improves supervised baselines and is a competitive alternative to the PPO RLHF implementation of earlier work while being simpler, easier to tune and cheaper to run (arXiv:2305.10425, v1 2023-05-17). It is the parallel work that complicates any "DPO invented the PPO-free route" story, because DPO appeared 12 days later with a different derivation (B05-35), and a text search of DPO's first arXiv version finds no citation of SLiC-HF, so they look like parallel efforts (Inference). Sources: Zhao et al.