Self-Instruct, where a language model writes its own instruction data from 175 seed tasks
Self-Instruct starts from 175 human-written seed tasks, has a language model generate new instructions and examples, filters them, and fine-tunes the same model on the result, which gives instruction…
- Date
- 20 December 2022
- Who
- University of Washington, AI2, Arizona State, JHU, Tehran Polytechnic
- People
- Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Landmark · Significance: 4/5 · Org(s): University of Washington, AI2, Arizona State, JHU, Tehran Polytechnic · People: Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi · Confidence: High Primary sources: arXiv:2212.10560 (v1 2022-12-20; ACL 2023) · parallel work, Unnatural Instructions, arXiv:2212.09689 (v1 2022-12-19)
One-liner. Self-Instruct starts from 175 human-written seed tasks, has a language model generate new instructions and examples, filters them, and fine-tunes the same model on the result, which gives instruction data without a labeling workforce.
Why it happened. By late 2022 the best instruction followers depended on human-written data that is limited in quantity, diversity and creativity, in the authors' words. That data was either public academic collections such as Super-NaturalInstructions (B05-07, B05-08) or OpenAI's private customer prompts and paid labelers (B05-13). The question was whether that bottleneck could be bypassed with the model's own generations. Three of the authors (Mishra, Khashabi, Hajishirzi) came from the Natural Instructions line (B05-07) (paper).
The idea. Bootstrap by using a small human seed set to prompt a model for new tasks, keeping the good ones, adding them to the pool, repeating, and then fine-tuning on the pool.
How it works. (1) A pool starts with 175 seed tasks (one instruction and one instance each). (2) At each step, 8 tasks are sampled from the pool as in-context examples (6 human-written, 2 model-generated) and the model writes new instructions. (3) The model classifies whether each is a classification task, then generates input-output instances. (4) A filter keeps a new instruction only if its ROUGE-L similarity with every existing one is below 0.7, and drops items with keywords the model cannot handle (such as image or graph). The result was over 52K instructions and about 82K instances, generated with vanilla GPT-3 (davinci, 175B) for around $600 of API cost at December 2022 prices, and GPT-3 was then fine-tuned on it through OpenAI's fine-tuning API (paper, App. A).
Results. On Super-NaturalInstructions, the method gave a 33% absolute improvement over vanilla GPT-3, which the authors describe as on par with InstructGPT-001. On a separate set of expert-written instructions for novel uses, human evaluation put Self-Instruct far above models tuned on existing public instruction datasets but still left a 5% absolute gap behind InstructGPT-001 (so the "on par" claim holds for the academic benchmark and does not extend to novel real-world instructions). Lab-measured; ACL 2023.
How it spread.
| Project | What it did | Relationship | Date | Lag vs. Self-Instruct |
|---|---|---|---|---|
| Unnatural Instructions (Honovich, Scialom, Levy, Schick) | Model given three seed examples to produce 64,000 examples, expanded by rephrasing to about 240,000; reported to rival manually curated datasets and to surpass T0++ and Tk-Instruct | parallel, one day earlier | 2022-12-19 | -1 day |
| Stanford Alpaca | Reran the recipe with text-davinci-003 as generator on LLaMA 7B, under $600 in total (B05-28) | derived | 2023-03-13 | 12 weeks |
| WizardLM (Evol-Instruct) | Rewrites instructions step by step into harder ones using an LLM (B05-32a) | derived, extension | 2023-04-24 | 18 weeks |
| Orca (Microsoft) | Learns from GPT-4 explanation traces instead of bare answers (B05-32a) | related (imitation line) | 2023-06-05 | 24 weeks |
| Dolly 2.0 (Databricks) | Deliberately used human-written data instead, for licensing reasons (B05-31) | alternative route | 2023-04-12 | 16 weeks |
Diffusion was fast because the paper, the data and the code were public and the generator was an API call.
Why it mattered. It is the template for cheap imitation, because it substitutes model-written demonstrations for human ones, which moved the open-source conversation from "who can afford labelers" to "which model do you distill", and it set up the argument over imitation versus capability (B05-34). It does not use preferences or RL.
Nuance, controversy and myths. (1) The generator matters. Self-Instruct used vanilla GPT-3, so the outputs carried no RLHF style, while Alpaca used text-davinci-003, itself RLHF-trained, so Alpaca inherits an RLHF-shaped style by distillation (B05-19, B05-28). (2) "Almost annotation-free" still needs 175 human seeds and a strong generator someone paid to build. (3) Terms of service of proprietary generators became a legal question for derived models.
Interview kit.
- 30-second version: Self-Instruct (UW and AI2, December 2022) bootstraps instruction data from 175 seed tasks using the model itself, filters near-duplicates and fine-tunes on the result; it matched InstructGPT-001 on an academic benchmark for about $600 and is the recipe behind Alpaca.
- Likely follow-ups: Is it RLHF? → No; it is supervised fine-tuning on model-written data. Did it match InstructGPT? → On Super-NaturalInstructions yes; on novel expert-written instructions it left a 5-point gap. Was it first? → Unnatural Instructions appeared one day earlier, so call it one of two near-simultaneous efforts.
- Common mistake: Saying Self-Instruct needs a strong aligned model; the original used vanilla GPT-3.
- Connect it to: B05-07, B05-28, B05-32a, B04.
Sources. 1. Wang et al., arXiv:2212.10560 · 2. Honovich et al., arXiv:2212.09689