Google's FLAN, BigScience's T0 and Flan-T5/PaLM instruction-tune models on academic tasks
Fine-tune a big pretrained model on many tasks written as natural-language instructions and it follows instructions on tasks it never saw, with no human preference data at all.
- Date
- 3 September 2021
- Who
- Google Research; BigScience/Hugging Face
- People
- Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, Quoc V. Le (FLAN); Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach and 35+ others (T0); Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph (Flan-PaLM)
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Landmark · Significance: 4/5 · Org(s): Google Research; BigScience/Hugging Face · People: Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, Quoc V. Le (FLAN); Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach and 35+ others (T0); Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph (Flan-PaLM) · Confidence: High Primary sources: FLAN, arXiv:2109.01652 (v1 2021-09-03) · T0, arXiv:2110.08207 (v1 2021-10-15) · Flan-PaLM/Flan-T5, arXiv:2210.11416 (v1 2022-10-20) · Flan Collection, arXiv:2301.13688 (v1 2023-01-31)
One-liner. Fine-tune a big pretrained model on many tasks written as natural-language instructions and it follows instructions on tasks it never saw, with no human preference data at all.
Why it happened. GPT-3 was good at few-shot prompting but weak zero-shot, particularly on reading comprehension and natural language inference; the FLAN authors guessed that without example prompts the input did not resemble pretraining text (FLAN §1). The bet was that if existing supervised NLP datasets were verbalized as instructions, the model would learn the skill of following instructions as well as the tasks. T0 tested the same idea from the open-science side, asking whether explicit multitask training can induce the zero-shot generalization that GPT-3 seemed to get implicitly from pretraining (T0 abstract).
The idea. "Instruction tuning" is supervised fine-tuning on a mixture of datasets, each rewritten with several instruction templates, and evaluated on held-out task types.
How it works. FLAN took a 137B-parameter pretrained language model (the LaMDA-PT base), instruction-tuned it on over 60 NLP datasets grouped into task clusters (each cluster held out in turn for evaluation), with ten hand-written templates per dataset, some deliberately "turned around" (e.g., generate a movie review given a sentiment). Instruction tuning took about 60 hours on a 128-core TPUv3 (FLAN paper). T0 did the same with T5-11B on prompted datasets from the PromptSource collection. Flan-PaLM later scaled to 1.8K tasks, PaLM 540B and added chain-of-thought data (Chung et al.).
Results. FLAN 137B beat zero-shot GPT-3 175B on 20 of 25 datasets and few-shot GPT-3 on ANLI, RTE, BoolQ, AI2-ARC, OpenbookQA and StoryCloze. Instruction tuning hurt held-out performance for models of 8B parameters or fewer, i.e. it only worked at scale. T0 outperformed models up to 16x its size on several held-out tasks. Flan-PaLM 540B reached 72.2% on five-shot MMLU against 69.3% for PaLM 540B (75.2% with chain-of-thought plus self-consistency, the figure in the paper's abstract); its +9.4% is the normalized average gain over PaLM 540B across the evaluation suites (MMLU, BBH, TyDiQA, MGSM) and does not apply to MMLU alone. The Flan-T5 checkpoints were released openly (Chung et al., Table 1). All lab-reported on academic benchmarks.
How it spread.
| Lab/project | Response | Relationship | Date | Lag vs. FLAN |
|---|---|---|---|---|
| BigScience / Hugging Face | T0 (open, T5-11B) | parallel, independent (same question, open-science side) | 2021-10-15 | 6 weeks |
| OpenAI | Cites FLAN and T0 and compares in InstructGPT, where labelers prefer InstructGPT to FLAN/T0-tuned GPT-3 | benchmarked against | 2022-01-27 | 5 months |
| Flan-PaLM and open Flan-T5 | same lab, scale-up | 2022-10-20 | 13 months | |
| BigScience | BLOOMZ and mT0, multilingual multitask fine-tuning (arXiv:2211.01786) | derived (T0 recipe) | 2022-11-03 | 14 months |
| Meta | OPT-IML, instruction meta-learning on OPT (arXiv:2212.12017) | derived | 2022-12-22 | 16 months |
| Meta | Llama 2 SFT bootstraps from public instruction data (Flan) | adopted the data | 2023-07-18 | 22 months |
| Open-source | Alpaca/Vicuna-style models use model-generated instruction data instead (B05-23, B05-28) | alternative route | 2023-03 | 18 months |
Because FLAN and T0 used public datasets and (for T5/Flan-T5) public weights, diffusion was fast and cheap, and the constraints were model scale and licensing.
Why it mattered. It established supervised instruction tuning as a standard stage between pretraining and anything fancier. Every later pipeline (B05-13, B05-37, B06) starts with some form of SFT on instruction data.
Nuance, controversy and myths. (1) "Instruction tuning" is not RLHF. FLAN/T0 use academic NLP datasets and no human preferences; InstructGPT used customer-style prompts plus human demonstrations and rankings. (2) OpenAI's own test found the academic route insufficient for an assistant. On its API prompt distribution, FLAN- and T0-tuned GPT-3 were preferred to the SFT baseline only 29.8% and 26.8% of the time versus 73.4% for InstructGPT, which OpenAI read as academic tasks not matching real usage (InstructGPT §1). (3) FLAN's 137B base was LaMDA-PT, and its headline comparison is to base GPT-3, not to OpenAI's own instruction-following API models, so "FLAN beat GPT-3" compares different base models and different training.
Interview kit.
- 30-second version: Google (FLAN) and BigScience (T0) showed in 2021 that fine-tuning on lots of tasks phrased as instructions makes models follow new instructions zero-shot. That is supervised, not RLHF; OpenAI then showed customer-style data plus human preferences was better for assistants.
- Likely follow-ups: Is instruction tuning the same as RLHF? → No; it is SFT on instruction-formatted data; RLHF adds preference-based optimization. Why did it only work for big models? → Small models lacked spare capacity (FLAN's ablation).
- Common mistake: Crediting InstructGPT with inventing instruction tuning.
- Connect it to: B05-13, B05-23, B05-33, B01.
Sources. 1. Wei et al., FLAN · 2. Sanh et al., T0 · 3. Chung et al., Flan-PaLM/T5 · 4. Ouyang et al.