Ibarz et al. at DeepMind combine demonstrations and preferences to learn Atari rewards
The Ibarz et al. paper was the first empirical follow-up to the 2017 paper, from the DeepMind safety team with Amodei, and appeared four days before Leike's agenda paper.
- Date
- 15 November 2018
- Who
- DeepMind, OpenAI
- People
- Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, Dario Amodei
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): DeepMind, OpenAI · People: Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, Dario Amodei · Confidence: High The Ibarz et al. paper was the first empirical follow-up to the 2017 paper, from the DeepMind safety team with Amodei, and appeared four days before Leike's agenda paper. It combined two forms of human feedback, expert demonstrations and trajectory preferences, to train a reward model and a DQN-based agent on nine Atari games; the approach beat the imitation-learning baseline in seven games and reached strictly superhuman scores in two without using the game reward, and the authors report fitting-quality checks, reward-hacking problems and the effect of label noise (arXiv:1811.06521, v1 2018-11-15). It belongs here because demonstrations-then-preferences is the shape later assistants took (SFT on demonstrations, then a reward model from comparisons, B05-13), and it shows DeepMind running its own line of this work from the start, which matters for the "why was Google late" question (B05-12, B05-27a). Ziegler et al. cite it as work on "relatively simple simulated environments" (Ziegler et al.). Sources: Ibarz et al. · Ziegler et al.