Hugging Face's Zephyr-7B uses distilled DPO on AI-feedback data
Zephyr-7B is the earliest high-profile demonstration I found that direct preference optimization plus AI feedback could stand in for an RLHF pipeline in open post-training.
- Date
- 25 October 2023
- Who
- Hugging Face (the H4 team)
- People
- Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, Thomas Wolf
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): Hugging Face (the H4 team) · People: Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, Thomas Wolf · Confidence: High Zephyr-7B is the earliest high-profile demonstration I found that direct preference optimization plus AI feedback could stand in for an RLHF pipeline in open post-training. The recipe starts from Mistral-7B. It applies distilled SFT on UltraChat (1.47M multi-turn dialogues generated by a model), collects AI feedback by ranking an ensemble of model completions with GPT-4 (the UltraFeedback dataset), and then runs distilled DPO on those preferences, a few hours of training with no sampling during fine-tuning and, in the authors' words, no human annotation (arXiv:2310.16944, v1 2023-10-25). The authors report that Zephyr-7B beats Llama 2-Chat-70B, "the best open-access RLHF-based model", on MT-Bench (7.34 versus 6.86), while on AlpacaEval it scores 90.6% against Llama 2-Chat-70B's 92.7% in the paper's table, so the claim depends on the benchmark (Table 1). It came 21 weeks after DPO's first version (B05-35) and 14 weeks after Llama 2-Chat (B05-37), and combines two threads of this chapter, distillation from a stronger model (B05-28) and AI feedback (B05-21, B05-38a). Two caveats apply. Both benchmarks use model judges, whose biases the Arena paper documents (B05-32b). The preference labels also come from GPT-4, a proprietary model that was itself trained with human feedback (B05-29), so "no human annotation" describes this pipeline and says nothing about the lineage behind it. Depth is in B06. Sources: Tunstall et al.