UC Berkeley's "The False Promise of Imitating Proprietary LLMs" finds that imitation models close little of the gap

The Berkeley authors fine-tuned a series of models (1.5B-13B parameters; 0.3M-150M tokens of imitation data) on ChatGPT outputs, as in Alpaca and Self-Instruct.

Date
25 May 2023
Who
UC Berkeley
People
Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, Dawn Song
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): UC Berkeley · People: Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, Dawn Song · Confidence: High The Berkeley authors fine-tuned a series of models (1.5B-13B parameters; 0.3M-150M tokens of imitation data) on ChatGPT outputs, as in Alpaca and Self-Instruct. Crowd raters found the outputs competitive with ChatGPT, but targeted automatic evaluations showed that imitation models close little of the gap on tasks not heavily covered by the imitation data. The models copy ChatGPT's style without its factuality, and raters can miss this. Their conclusion is that bridging the gap requires stronger base models or an "unwieldy amount" of imitation data (arXiv:2305.15717, v1 2023-05-25). This is the empirical version of Schulman's April prediction (B05-32) and the counterweight to Alpaca (B05-28). Interview trap: the paper concerns small base models imitating ChatGPT, not distillation in general; later reasoning-model distillation worked well with stronger bases and verified data (B08). Sources: Gudibande et al.

Read it in the deep dive