Mishra et al. publish Natural Instructions, 61 tasks with human-written instructions
Natural Instructions is a large crowdsourced-instruction benchmark built to test whether a model can learn a new task by reading its instructions.
- Date
- 18 April 2021
- Who
- Allen Institute for AI, Arizona State, UW
- People
- Swaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh Hajishirzi
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): Allen Institute for AI, Arizona State, UW · People: Swaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh Hajishirzi · Confidence: High Natural Instructions is a large crowdsourced-instruction benchmark built to test whether a model can learn a new task by reading its instructions. It has 61 tasks, their human-written instructions (taken from the crowdsourcing instructions used to create existing NLP datasets) and 193k instances, and models trained on some tasks were tested on unseen ones (arXiv:2104.08773, v1 2021-04-18).
It was not the first attempt. The paper itself cites Weller et al.'s ZEST (task descriptions as questions, arXiv:2011.08115, 2020-11-16) and Efrat and Levy (2020), and Zhong et al.'s meta-tuning, the same idea at smaller scale, appeared eight days earlier (arXiv:2104.04670, v1 2021-04-10). The authors claim only that, to their knowledge, it was the first work to show the benefit of instructions for cross-task generalization. Its successor Super-NaturalInstructions (2022-04-16) scaled this to 1,616 tasks across 76 task types (arXiv:2204.07705).
It matters here for two reasons. It is the academic root of "instructions in, behavior out" that FLAN and T0 scaled (B05-08), and its authors (Mishra, Khashabi, Hajishirzi) reappear on Self-Instruct, which produced the data behind Alpaca (B05-23, B05-28). Sources: Mishra et al. · Wang et al. · Weller et al. · Zhong et al.