AlpacaFarm, a Stanford simulator that cuts the cost of RLHF experiments from about $3,150 to about $70
AlpacaFarm, an open simulator from Stanford that follows Alpaca (B05-28), replaces crowdworkers with LLM-simulated annotators (prompts designed to mimic human variability and noise), supplies an…
- Date
- 22 May 2023
- Who
- Stanford (with collaborators)
- People
- Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, Tatsunori Hashimoto
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): Stanford (with collaborators) · People: Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, Tatsunori Hashimoto · Confidence: High AlpacaFarm, an open simulator from Stanford that follows Alpaca (B05-28), replaces crowdworkers with LLM-simulated annotators (prompts designed to mimic human variability and noise), supplies an automatic evaluation validated against real human judgments, and ships reference implementations of PPO, DPO, best-of-n, expert iteration and other methods. The authors report that the simulated annotator agrees with a majority of three humans about as well as a held-out human does (65% versus 66%) while costing 25 times less ($300 to $12 per 1,000 examples), that a full experiment costs about $70 against $3,150 with real human feedback, that method rankings in the simulator match rankings from training on 10,000 pairs of real human feedback (Spearman 0.98), that it reproduces reward-model over-optimization (B05-18), and that their reference PPO implementation yields a +10% win-rate improvement against text-davinci-003 (the abstract's wording) (arXiv:2305.14387, v1 2023-05-22). It mattered because it took the cost of an RLHF-method experiment from thousands of dollars to tens, in the authors' accounting, which plausibly helped open work move from Alpaca to DPO and Zephyr (Inference; B05-39a). The caveat is that simulated feedback inherits the biases of the simulating model. Sources: Dubois et al.