Direct Preference Optimization

DPO reparameterizes the RLHF objective so the optimal policy can be extracted in closed form, letting preference pairs be fit with a simple classification loss, with no separate reward model, no…

Date
29 May 2023
Who
Stanford
People
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher Manning, Chelsea Finn
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): Stanford · People: Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher Manning, Chelsea Finn · Confidence: High DPO reparameterizes the RLHF objective so the optimal policy can be extracted in closed form, letting preference pairs be fit with a simple classification loss, with no separate reward model, no sampling during fine-tuning and little hyperparameter tuning (arXiv:2305.18290, v1 2023-05-29). The authors report that it exceeds PPO-based RLHF at controlling sentiment and matches or improves response quality in summarization and single-turn dialogue, while being substantially simpler to implement and train (abstract). It removed the PPO engineering barrier of B05-03. Precedents and parallel work: Google's SLiC-HF, published 12 days earlier with a different derivation and the same pitch of a simpler alternative to PPO (B05-32d); DPO's first version does not cite it. Adoption: the first high-profile open use I found is Hugging Face's Zephyr-7B, 21 weeks later (B05-39a), and DeepSeek LLM Chat used SFT then DPO in place of PPO (B05-30a); Llama 2-Chat, released seven weeks after DPO's first version, still used PPO (B05-37). Whether DPO became the default for open post-training, the variants it spawned and the debate on whether it matches online RL are in B06. Sources: Rafailov et al. · Zhao et al., SLiC-HF

Read it in the deep dive