DeepMind trains GopherCite, its first large-model RLHF

DeepMind's GopherCite paper, "Teaching language models to support answers with verified quotes", trained a 280-billion-parameter Gopher with what the authors call reinforcement learning from human…

Date
21 March 2022
Who
DeepMind
People
Jacob Menick, Maja Trębacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, Nat McAleese
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): DeepMind · People: Jacob Menick, Maja Trębacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, Nat McAleese · Confidence: High DeepMind's GopherCite paper, "Teaching language models to support answers with verified quotes", trained a 280-billion-parameter Gopher with what the authors call reinforcement learning from human preferences (RLHP) to answer open-book questions while quoting evidence, and to abstain when unsure. Humans rated 80% of its answers high quality on a NaturalQuestions subset and 67% on an ELI5 subset; abstaining on the third of questions it was least sure about raised those to 90% and 80%, and the authors stress that "supported by evidence" is not the same as true (TruthfulQA analysis) (arXiv:2203.11147, v1 2022-03-21). It matters here because it was DeepMind's own large-language-model RLHF, published two months after OpenAI's InstructGPT blog and six months before Sparrow (B05-16), so Sparrow was not DeepMind's first, and Hugging Face's December 2022 RLHF explainer lists Gopher and GopherCite among RLHF language models (HF blog). It sharpens the "why was Google late" story, since Alphabet had RLHF in DeepMind while LaMDA and PaLM lived in Google Brain (B05-12, B05-27a). Sources: Menick et al. · Hugging Face, 2022-12-09

Read it in the deep dive