OpenAI's weak-to-strong generalization paper
The paper is an empirical follow-up to the Superalignment announcement (B05-36) and is co-authored by that program's two announced leads.
- Date
- 14 December 2023
- Who
- OpenAI
- People
- Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, Jeff Wu
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · People: Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, Jeff Wu · Confidence: High The paper is an empirical follow-up to the Superalignment announcement (B05-36) and is co-authored by that program's two announced leads. The abstract's setup is an analogy for supervising superhuman models, and it asks whether a weak model's labels can elicit the full capabilities of a much stronger one. On NLP tasks, chess puzzles and reward modeling with GPT-4-family models, naive fine-tuning of strong models on weak-model labels made them consistently better than their weak supervisors ("weak-to-strong generalization") but recovered far from their full capabilities, which the authors read as a sign that techniques like RLHF may scale poorly to superhuman models without further work; simple methods helped, for example fine-tuning GPT-4 with a GPT-2-level supervisor plus an auxiliary confidence loss recovered close to GPT-3.5-level performance on NLP tasks (arXiv:2312.09390, v1 2023-12-14). It is the experimental counterpart to the claim in B05-36, and it sits beside Anthropic's scalable-oversight proof of concept (B05-18a) and OpenAI's critic work (B05-14a, B05-40c). Depth is in B22. Sources: Burns et al.