Paul Christiano's first-hand account of how RLHF research began
In "Thoughts on the impact of RLHF research" (Alignment Forum, dated 2023-01-25 on the page), Christiano says work on this direction began in 2015; that after joining OpenAI full time in 2017,…
- Date
- 25 January 2023
- Who
- OpenAI, DeepMind, Alignment Research Center
- People
- Paul Christiano, Dario Amodei, Geoffrey Irving, Jan Leike
- Confidence
- High (first-hand and dated), with the caveat that it is one participant's retros
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): OpenAI, DeepMind, Alignment Research Center · People: Paul Christiano, Dario Amodei, Geoffrey Irving, Jan Leike · Confidence: High (first-hand and dated), with the caveat that it is one participant's retrospective and an argument for his own side In "Thoughts on the impact of RLHF research" (Alignment Forum, dated 2023-01-25 on the page), Christiano says work on this direction began in 2015; that after joining OpenAI full time in 2017, Christiano was pushed by manager Dario Amodei to use real human feedback instead of synthetic labels in the first project; that Geoffrey Irving joined in mid-2017, championed applying it to language models and led a larger team that finished the first language-model work in 2019; that after Irving left for DeepMind Christiano led the team through follow-up papers aimed at production use; and that Christiano left in early 2021, with Jan Leike taking over (post). He argues the research was net positive for alignment. It matters because it is the best first-hand source on who proposed what, it closes much of the attribution backlog (Backlog), and it places DeepMind people (Irving, Leike) in the origin story, so RLHF cannot be treated as purely an OpenAI invention. Read it beside Christiano's later, more ambivalent remark that ChatGPT's acceleration of timelines was probably net negative (Dwarkesh, 2023-10-31; B05-02, B05-20). The two statements are nine months apart and concern different things (the research versus the product's press effect). I read the post through a fetch summary, not line by line. Sources: Christiano, Alignment Forum · Dwarkesh interview