Mishra et al. publish Natural Instructions, 61 tasks with human-written instructions

Natural Instructions is a large crowdsourced-instruction benchmark built to test whether a model can learn a new task by reading its instructions.

Date
18 April 2021
Who
Allen Institute for AI, Arizona State, UW
People
Swaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh Hajishirzi
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): Allen Institute for AI, Arizona State, UW · People: Swaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh Hajishirzi · Confidence: High Natural Instructions is a large crowdsourced-instruction benchmark built to test whether a model can learn a new task by reading its instructions. It has 61 tasks, their human-written instructions (taken from the crowdsourcing instructions used to create existing NLP datasets) and 193k instances, and models trained on some tasks were tested on unseen ones (arXiv:2104.08773, v1 2021-04-18).

It was not the first attempt. The paper itself cites Weller et al.'s ZEST (task descriptions as questions, arXiv:2011.08115, 2020-11-16) and Efrat and Levy (2020), and Zhong et al.'s meta-tuning, the same idea at smaller scale, appeared eight days earlier (arXiv:2104.04670, v1 2021-04-10). The authors claim only that, to their knowledge, it was the first work to show the benefit of instructions for cross-task generalization. Its successor Super-NaturalInstructions (2022-04-16) scaled this to 1,616 tasks across 76 task types (arXiv:2204.07705).

It matters here for two reasons. It is the academic root of "instructions in, behavior out" that FLAN and T0 scaled (B05-08), and its authors (Mishra, Khashabi, Hajishirzi) reappear on Self-Instruct, which produced the data behind Alpaca (B05-23, B05-28). Sources: Mishra et al. · Wang et al. · Weller et al. · Zhong et al.

Read it in the deep dive