Direct Preference Optimization
DPO reparameterizes the RLHF objective so the optimal policy can be extracted in closed form, letting preference pairs be fit with a simple classification loss, with no separate reward model, no…
- Date
- 29 May 2023
- Who
- Stanford
- People
- Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher Manning, Chelsea Finn
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): Stanford · People: Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher Manning, Chelsea Finn · Confidence: High DPO reparameterizes the RLHF objective so the optimal policy can be extracted in closed form, letting preference pairs be fit with a simple classification loss, with no separate reward model, no sampling during fine-tuning and little hyperparameter tuning (arXiv:2305.18290, v1 2023-05-29). The authors report that it exceeds PPO-based RLHF at controlling sentiment and matches or improves response quality in summarization and single-turn dialogue, while being substantially simpler to implement and train (abstract). It removed the PPO engineering barrier of B05-03. Precedents and parallel work: Google's SLiC-HF, published 12 days earlier with a different derivation and the same pitch of a simpler alternative to PPO (B05-32d); DPO's first version does not cite it. Adoption: the first high-profile open use I found is Hugging Face's Zephyr-7B, 21 weeks later (B05-39a), and DeepSeek LLM Chat used SFT then DPO in place of PPO (B05-30a); Llama 2-Chat, released seven weeks after DPO's first version, still used PPO (B05-37). Whether DPO became the default for open post-training, the variants it spawned and the debate on whether it matches online RL are in B06. Sources: Rafailov et al. · Zhao et al., SLiC-HF