OpenAI's Superalignment announcement says current alignment techniques "will not scale to superintelligence"

OpenAI announced a Superalignment team co-led by Sutskever and Leike with 20% of compute secured to date dedicated to the effort over four years.

Date
5 July 2023
Who
OpenAI
People
Ilya Sutskever, Jan Leike
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · People: Ilya Sutskever, Jan Leike · Confidence: High OpenAI announced a Superalignment team co-led by Sutskever and Leike with 20% of compute secured to date dedicated to the effort over four years. The post names RLHF as an example of current alignment techniques that rely on humans' ability to supervise AI, says humans will not be able to reliably supervise much smarter systems, and concludes that current alignment techniques will not scale to superintelligence (RLHF is named as an example, not singled out); the goal was a roughly human-level automated alignment researcher (OpenAI, 2023-07-05). That is OpenAI, with a co-author of the 2017 paper (Leike) as co-lead, stating RLHF's ceiling, in the framing of B05-04; the same position (RLHF as a building block, not a sufficient solution) was already in OpenAI's 2022-08-24 alignment-approach post (OpenAI). OpenAI disbanded the team in May 2024, days after Sutskever and Leike announced their departures (CNBC, 2024-05-17); Leike joined Anthropic (CNBC, 2024-05-28). An empirical paper co-authored by Leike and Sutskever, on weak-to-strong generalization, followed in December (B05-40a); later work on scalable oversight is in B22. Sources: OpenAI · CNBC

Read it in the deep dive