LIMA, "Less Is More for Alignment", and its test of the superficial alignment hypothesis with 1,000 examples
A 65B LLaMA model fine-tuned on only 1,000 carefully chosen examples, with no RL and no preference modeling, was preferred or tied with GPT-4 in 43% of cases, and the authors take that as support for…
- Date
- 18 May 2023
- Who
- Meta AI, Carnegie Mellon, USC, Tel Aviv University
- People
- Chunting Zhou, Pengfei Liu, Omer Levy, Luke Zettlemoyer and others
- Confidence
- High (setup); Medium (generalization of the claim)
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Landmark · Significance: 4/5 · Org(s): Meta AI, Carnegie Mellon, USC, Tel Aviv University · People: Chunting Zhou, Pengfei Liu, Omer Levy, Luke Zettlemoyer and others · Confidence: High (setup); Medium (generalization of the claim) Primary sources: arXiv:2305.11206 (v1 2023-05-18)
One-liner. A 65B LLaMA model fine-tuned on only 1,000 carefully chosen examples, with no RL and no preference modeling, was preferred or tied with GPT-4 in 43% of cases, and the authors take that as support for the idea that alignment mostly teaches style and format.
Why it happened. After InstructGPT and ChatGPT, open models were being fine-tuned on tens of thousands (Alpaca at 52K) or millions of examples. The LIMA authors asked how much of what makes a model useful comes from pretraining versus from instruction tuning and RL.
The idea. The Superficial Alignment Hypothesis says that a model's knowledge and capabilities are learned almost entirely during pretraining, and that alignment teaches only which sub-distribution of formats to use when interacting with users, so a small clean set can be enough (paper §1).
How it works. The training set is 1,000 prompt-response pairs, about 750,000 tokens, of which 750 were selected from community forums (Stack Exchange STEM and other, wikiHow, Reddit's r/WritingPrompts) and 250 were written or adapted by the authors for a uniform assistant style. LLaMA 65B was fine-tuned on them with supervised fine-tuning and the standard loss. Evaluation used 300 held-out challenging test prompts, rated by humans and GPT-4.
Results. In a controlled human study, LIMA's responses were equivalent to or preferred over GPT-4's in 43% of cases, over Bard's in 58%, and over text-davinci-003, "trained with human feedback", in 65% (abstract). Lab-measured on a small prompt set.
How it spread. Meta's own Llama 2 paper (2023-07-18, 2 months later) cites it for the same finding. Meta collected only 27,540 high-quality vendor annotations, found them better than millions of third-party examples, and describes this as similar in spirit to LIMA (Llama 2 §3.1). Later data-curation work is in B06 and B04.
Why it mattered. It reframed the central question. If capability is already in the base model, alignment data is a specification problem and scale matters less. It undercut the assumption that RLHF's gains come from the volume of human labels.
Nuance, controversy and myths. (1) LIMA does not show RLHF is unnecessary, since Llama 2's authors found that after SFT, reward-model-based RLHF improved results and argued the model's ability to write beyond annotators' skill was driven by RLHF (B05-37). (2) The test set is small, preferences are single-turn, and safety is thin. (3) LIMA's own data came from the web, including answers from community sites that were themselves often written by humans, so it is not a purely synthetic result. (4) Inference: GPT-4's large TruthfulQA gain and calibration loss after RLHF (B05-29) suggest RLHF changes more than surface style.
Interview kit.
- 30-second version: LIMA (Meta et al., May 2023) fine-tuned a 65B model on 1,000 examples and was judged equivalent to or better than GPT-4 in 43% of comparisons (so GPT-4 was strictly preferred in the other 57%), arguing most knowledge comes from pretraining and alignment is mostly style.
- Likely follow-ups: So is RLHF useless? → No; later work shows RL on preferences still helps, and reasoning-style RL changes capability (B08). How is that consistent with InstructGPT? → InstructGPT also said RLHF brings out abilities already in GPT-3 (B05-13).
- Common mistake: Reading "less is more" as "data quality doesn't matter at scale" or "no RL ever needed".
- Connect it to: B05-13, B05-25, B05-37, B04.
Sources. 1. Zhou et al., arXiv:2305.11206 · 2. Touvron et al., arXiv:2307.09288