Google's SLiC-HF, a simpler alternative to PPO published twelve days before DPO

Sequence Likelihood Calibration with Human Feedback (SLiC-HF) adapted Google's earlier SLiC method to learn from human preference pairs, including preference data collected for a different model (an…

Date
17 May 2023
Who
Google (the paper lists Google DeepMind and Google Research)
People
Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, Peter J. Liu
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): Google (the paper lists Google DeepMind and Google Research) · People: Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, Peter J. Liu · Confidence: High Sequence Likelihood Calibration with Human Feedback (SLiC-HF) adapted Google's earlier SLiC method to learn from human preference pairs, including preference data collected for a different model (an off-policy, offline setting), and on TL;DR summarization the authors report that it significantly improves supervised baselines and is a competitive alternative to the PPO RLHF implementation of earlier work while being simpler, easier to tune and cheaper to run (arXiv:2305.10425, v1 2023-05-17). It is the parallel work that complicates any "DPO invented the PPO-free route" story, because DPO appeared 12 days later with a different derivation (B05-35), and a text search of DPO's first arXiv version finds no citation of SLiC-HF, so they look like parallel efforts (Inference). Sources: Zhao et al.

Read it in the deep dive