DeepMind publishes Sparrow, an RLHF agent trained with rules and evidence
Sparrow was an information-seeking dialogue agent built from a dialogue-prompted Chinchilla 70B and trained with RLHF using two additions, a set of natural-language rules that raters judged…
- Date
- 22 September 2022
- Who
- DeepMind
- People
- Amelia Glaese, Nat McAleese, Maja Trębacz and others; Geoffrey Irving (senior)
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): DeepMind · People: Amelia Glaese, Nat McAleese, Maja Trębacz and others; Geoffrey Irving (senior) · Confidence: High Sparrow was an information-seeking dialogue agent built from a dialogue-prompted Chinchilla 70B and trained with RLHF using two additions, a set of natural-language rules that raters judged separately (so reward models could be rule-conditional) and evidence from search shown to raters for factual claims. It was preferred to baselines more often and broke the rules in 8% of adversarial probing; the evidence supported answers 78% of the time (blog 2022-09-22; arXiv:2209.14375, v1 2022-09-28). DeepMind did not release it; in January 2023 Demis Hassabis said it was considering a "private beta" in 2023 and was delaying to add RL-based features such as source citation (TIME, 2023-01-12).
A Verge investigation found one annotator on a Remotasks project code-named "Dolphin", later identified as Sparrow, who was paid about $14 an hour to chat with it all day (Dzieza, 2023-06-20). Its senior author, Geoffrey Irving, was an author of the 2018 debate paper and the 2019 GPT-2 paper while at OpenAI, so the lineage runs from the OpenAI/DeepMind safety collaboration into DeepMind's own RLHF (B05-04, B05-05). Sources: DeepMind blog · Glaese et al. · TIME