AlpacaFarm, a Stanford simulator that cuts the cost of RLHF experiments from about $3,150 to about $70

AlpacaFarm, an open simulator from Stanford that follows Alpaca (B05-28), replaces crowdworkers with LLM-simulated annotators (prompts designed to mimic human variability and noise), supplies an…

Date
22 May 2023
Who
Stanford (with collaborators)
People
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, Tatsunori Hashimoto
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): Stanford (with collaborators) · People: Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, Tatsunori Hashimoto · Confidence: High AlpacaFarm, an open simulator from Stanford that follows Alpaca (B05-28), replaces crowdworkers with LLM-simulated annotators (prompts designed to mimic human variability and noise), supplies an automatic evaluation validated against real human judgments, and ships reference implementations of PPO, DPO, best-of-n, expert iteration and other methods. The authors report that the simulated annotator agrees with a majority of three humans about as well as a held-out human does (65% versus 66%) while costing 25 times less ($300 to $12 per 1,000 examples), that a full experiment costs about $70 against $3,150 with real human feedback, that method rankings in the simulator match rankings from training on 10,000 pairs of real human feedback (Spearman 0.98), that it reproduces reward-model over-optimization (B05-18), and that their reference PPO implementation yields a +10% win-rate improvement against text-davinci-003 (the abstract's wording) (arXiv:2305.14387, v1 2023-05-22). It mattered because it took the cost of an RLHF-method experiment from thousands of dollars to tens, in the authors' accounting, which plausibly helped open work move from Alpaca to DPO and Zephyr (Inference; B05-39a). The caveat is that simulated feedback inherits the biases of the simulating model. Sources: Dubois et al.

Read it in the deep dive