UC Berkeley's "The False Promise of Imitating Proprietary LLMs" finds that imitation models close little of the gap
The Berkeley authors fine-tuned a series of models (1.5B-13B parameters; 0.3M-150M tokens of imitation data) on ChatGPT outputs, as in Alpaca and Self-Instruct.
- Date
- 25 May 2023
- Who
- UC Berkeley
- People
- Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, Dawn Song
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): UC Berkeley · People: Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, Dawn Song · Confidence: High The Berkeley authors fine-tuned a series of models (1.5B-13B parameters; 0.3M-150M tokens of imitation data) on ChatGPT outputs, as in Alpaca and Self-Instruct. Crowd raters found the outputs competitive with ChatGPT, but targeted automatic evaluations showed that imitation models close little of the gap on tasks not heavily covered by the imitation data. The models copy ChatGPT's style without its factuality, and raters can miss this. Their conclusion is that bridging the gap requires stronger base models or an "unwieldy amount" of imitation data (arXiv:2305.15717, v1 2023-05-25). This is the empirical version of Schulman's April prediction (B05-32) and the counterweight to Alpaca (B05-28). Interview trap: the paper concerns small base models imitating ChatGPT, not distillation in general; later reasoning-model distillation worked well with stronger bases and verified data (B08). Sources: Gudibande et al.