WizardLM, Orca and Tulu, three papers that followed Alpaca

Three papers tested whether imitation can get past style, the question Alpaca left open (B05-28).

Date
24 April 2023
Who
academic and industry groups (see papers)
Confidence
High (paper claims); Medium (their evaluations, which are model- and small-sampl
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): academic and industry groups (see papers) · Confidence: High (paper claims); Medium (their evaluations, which are model- and small-sample-based) Three papers tested whether imitation can get past style, the question Alpaca left open (B05-28). WizardLM (Xu et al., arXiv v1 2023-04-24) proposed Evol-Instruct, in which an LLM rewrites instructions step by step into more complex ones, and fine-tuned LLaMA on the mix; the authors report human preference for Evol-Instruct instructions over human-made ones and, on the hardest portion, for WizardLM over ChatGPT (arXiv:2304.12244). Orca (Mukherjee et al., arXiv 2023-06-05) trained a 13B model on GPT-4 explanation traces, with ChatGPT as an intermediate teacher, arguing that earlier imitation models learned the teacher's style but not its reasoning; it reports more than 100% improvement over Vicuna-13B on Big-Bench Hard and parity with ChatGPT there, while trailing GPT-4 (arXiv:2306.02707). Tulu ("How Far Can Camels Go?", Wang et al., arXiv 2023-06-07) compared 12 open instruction datasets on models from 6.7B to 65B, found no single dataset best across skills, found that model- and human-preference evaluations did not reflect capability differences that benchmarks exposed, and reported that the best model reached on average 87% of ChatGPT and 73% of GPT-4 performance (arXiv:2306.04751). Together they are the step between Alpaca and the DPO-era open models (B05-39a), and Tulu's evaluation finding feeds the benchmark debate in B23 and the imitation critique in B05-34. Sources: WizardLM · Orca · Tulu

Read it in the deep dive