Google's RLAIF vs. RLHF paper tests AI preference labels against human ones
Google's paper is the first independent lab-scale test I found of the mechanism behind Constitutional AI (B05-21), which replaces human preference labels with labels from an off-the-shelf LLM.
- Date
- 1 September 2023
- Who
- People
- Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, Sushant Prakash
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): Google · People: Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, Sushant Prakash · Confidence: High Google's paper is the first independent lab-scale test I found of the mechanism behind Constitutional AI (B05-21), which replaces human preference labels with labels from an off-the-shelf LLM. Across summarization, helpful dialogue and harmless dialogue, the authors report that RLAIF and RLHF are preferred over a supervised baseline at similar rates (71% versus 73% for summarization, 63% versus 64% for helpful dialogue) and equally preferred when compared head to head on summarization and helpfulness, and that on harmlessness RLAIF scored higher (88% harmless versus 76% for RLHF and 64% for the SFT baseline). They also report that RLAIF can improve on SFT even when the labeler is the same size as the policy or the very same checkpoint, and introduce "direct RLAIF", which skips reward-model training and takes rewards straight from an LLM during RL, reported as better than the canonical version (arXiv:2309.00267, v1 2023-09-01, revised through v3 on 2024-09-03). It came 37 weeks after Constitutional AI, from a lab that had its own RLHF line at DeepMind (B05-13a). Depth on RLAIF is in B06. Sources: Lee et al.