Hugging Face's Zephyr-7B uses distilled DPO on AI-feedback data

Zephyr-7B is the earliest high-profile demonstration I found that direct preference optimization plus AI feedback could stand in for an RLHF pipeline in open post-training.

Date
25 October 2023
Who
Hugging Face (the H4 team)
People
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, Thomas Wolf
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): Hugging Face (the H4 team) · People: Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, Thomas Wolf · Confidence: High Zephyr-7B is the earliest high-profile demonstration I found that direct preference optimization plus AI feedback could stand in for an RLHF pipeline in open post-training. The recipe starts from Mistral-7B. It applies distilled SFT on UltraChat (1.47M multi-turn dialogues generated by a model), collects AI feedback by ranking an ensemble of model completions with GPT-4 (the UltraFeedback dataset), and then runs distilled DPO on those preferences, a few hours of training with no sampling during fine-tuning and, in the authors' words, no human annotation (arXiv:2310.16944, v1 2023-10-25). The authors report that Zephyr-7B beats Llama 2-Chat-70B, "the best open-access RLHF-based model", on MT-Bench (7.34 versus 6.86), while on AlpacaEval it scores 90.6% against Llama 2-Chat-70B's 92.7% in the paper's table, so the claim depends on the benchmark (Table 1). It came 21 weeks after DPO's first version (B05-35) and 14 weeks after Llama 2-Chat (B05-37), and combines two threads of this chapter, distillation from a stronger model (B05-28) and AI feedback (B05-21, B05-38a). Two caveats apply. Both benchmarks use model judges, whose biases the Arena paper documents (B05-32b). The preference labels also come from GPT-4, a proprietary model that was itself trained with human feedback (B05-29), so "no human annotation" describes this pipeline and says nothing about the lineage behind it. Depth is in B06. Sources: Tunstall et al.

Read it in the deep dive