Alpaca and Vicuna, LLaMA models fine-tuned on OpenAI model outputs for under $600 and about $300

Stanford fine-tuned LLaMA 7B on 52K instruction-response pairs generated by text-davinci-003 for under $600, and the result behaved like an assistant in casual tests.

Date
13 March 2023
Who
Stanford CRFM; LMSYS (UC Berkeley, CMU, Stanford, UCSD); Meta (LLaMA)
People
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, Tatsunori Hashimoto
Confidence
High (numbers from the authors); Medium (quality claims)
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Landmark · Significance: 4/5 · Org(s): Stanford CRFM; LMSYS (UC Berkeley, CMU, Stanford, UCSD); Meta (LLaMA) · People: Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, Tatsunori Hashimoto · Confidence: High (numbers from the authors); Medium (quality claims) Primary sources: Alpaca blog, 2023-03-13 · Vicuna blog, 2023-03-30

One-liner. Stanford fine-tuned LLaMA 7B on 52K instruction-response pairs generated by text-davinci-003 for under $600, and the result behaved like an assistant in casual tests.

Why it happened. Academia had no strong open instruction-following model. Meta released LLaMA to researchers on 2023-02-24 under a non-commercial license; the weights leaked on 4chan about a week later (The Verge, 2023-03-08). The Alpaca team combined that base with Self-Instruct (B05-23) and OpenAI's API as a data generator (my inference is that the API had become cheap enough to make this practical).

The idea. Replace human demonstrations and preferences with outputs from a stronger model.

How it works. Starting from the 175 human-written seed instruction-output pairs of Self-Instruct, prompt text-davinci-003 to generate more; the process produced 52K unique instructions with outputs and cost under $500 of API credit. Fine-tuning LLaMA 7B used Hugging Face tooling with FSDP and mixed precision and took 3 hours on eight 80GB A100s (under $100). In a blind pairwise comparison of 5 student authors on the Self-Instruct evaluation set, Alpaca won 90 comparisons to text-davinci-003's 89 (Alpaca). Vicuna-13B (LMSYS) fine-tuned LLaMA on about 70K user-shared ChatGPT conversations from ShareGPT for around $300, and rated itself above 90% of ChatGPT/Bard quality using GPT-4 as the judge, which its authors labeled "fun and non-scientific" (Vicuna).

Results. Company/lab-measured, tiny evaluations. Alpaca's authors state that it hallucinates (they cite a wrong capital of Tanzania) and shared it only for non-commercial research because LLaMA's license forbids commercial use, OpenAI's terms forbid using its outputs to build competing models, and the team had not built adequate safety measures. They took the public demo down because of hosting costs and weak content filters.

How it spread.

ProjectWhat it didRelationshipDateLag vs. Alpaca
Vicuna (LMSYS)ShareGPT conversations, multi-turnparallel, same base (LLaMA), different data2023-03-3017 days
Dolly 1.0 (Databricks)Instruction-tuned open model trained for under $30parallellate March 2023 (two weeks before Dolly 2.0)about 2 weeks
Dolly 2.0 (Databricks)Pythia-12B on 15,000 human-written prompt/response pairs by 5,000+ employees, commercially licensablealternative to Alpaca's licensing problem2023-04-1230 days
OpenAssistant (LAION)Volunteer-written conversation trees (B05-31)alternative to Alpaca's licensing problem2023-04-1432 days
WizardLM, Orca, TüluEvolved instructions, GPT-4 explanation traces, and a 12-dataset comparison (B05-32a)derived and critical follow-ups2023-04-24 to 2023-06-076 to 12 weeks
Llama 2-Chat (Meta)Industrial-grade SFT plus RLHF (B05-37)own product (an answer to the quality gap, not a derivative)2023-07-1818 weeks

Diffusion was instant because it needed only API credits, a base model and a single GPU server.

Why it mattered. It showed that the style of instruction following is cheap to copy by distillation, which lowered the perceived moat of RLHF and set off many open-weights releases (B19). It also set up the "distillation" argument later central in the DeepSeek discussion (B20).

Nuance, controversy and myths. (1) Alpaca is not RLHF. It is supervised fine-tuning on distilled outputs, though the generator, text-davinci-003, was RLHF-trained (B05-19). (2) Berkeley researchers argued that such imitation models mimic style but not factuality or capability (B05-34). (3) The licensing and ToS issue was open from day one. (4) Evaluations were tiny and used model judges.

Interview kit.

  • 30-second version: Alpaca (Stanford, 2023-03-13) fine-tuned LLaMA 7B on 52K instructions generated by text-davinci-003, for under $600, producing an assistant-like model. It's distillation with supervised fine-tuning, and no RLHF step.
  • Likely follow-ups: Why was it significant? → It proved cheap imitation of an aligned model's behavior. Why wasn't it a real ChatGPT replacement? → It copies style, hallucinates, and was licensed for research only.
  • Common mistake: Calling Alpaca an RLHF model.
  • Connect it to: B05-23, B05-33, B05-34, B19.

Sources. 1. Alpaca · 2. Vicuna · 3. The Verge · 4. Databricks, Dolly 2.0

Read it in the deep dive