OpenAI's weak-to-strong generalization paper

The paper is an empirical follow-up to the Superalignment announcement (B05-36) and is co-authored by that program's two announced leads.

Date
14 December 2023
Who
OpenAI
People
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, Jeff Wu
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · People: Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, Jeff Wu · Confidence: High The paper is an empirical follow-up to the Superalignment announcement (B05-36) and is co-authored by that program's two announced leads. The abstract's setup is an analogy for supervising superhuman models, and it asks whether a weak model's labels can elicit the full capabilities of a much stronger one. On NLP tasks, chess puzzles and reward modeling with GPT-4-family models, naive fine-tuning of strong models on weak-model labels made them consistently better than their weak supervisors ("weak-to-strong generalization") but recovered far from their full capabilities, which the authors read as a sign that techniques like RLHF may scale poorly to superhuman models without further work; simple methods helped, for example fine-tuning GPT-4 with a GPT-2-level supervisor plus an auxiliary confidence loss recovered close to GPT-3.5-level performance on NLP tasks (arXiv:2312.09390, v1 2023-12-14). It is the experimental counterpart to the claim in B05-36, and it sits beside Anthropic's scalable-oversight proof of concept (B05-18a) and OpenAI's critic work (B05-14a, B05-40c). Depth is in B22. Sources: Burns et al.

Read it in the deep dive