Two papers measure RLHF side effects, length bias and lost diversity
Two October 2023 papers quantified RLHF side effects that practitioners had suspected.
- Date
- 5 October 2023
- Who
- academic groups (see papers)
- Confidence
- High (the papers' own measurements)
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): academic groups (see papers) · Confidence: High (the papers' own measurements) Two October 2023 papers quantified RLHF side effects that practitioners had suspected. Singhal, Goyal, Xu and Durrett ("A Long Way to Go", arXiv v1 2023-10-05) found, in three settings, that reward improvements are largely driven by longer responses, that a purely length-based reward reproduces most of RLHF's downstream gains over supervised fine-tuning, and that the dominant source of the bias is reward models that are non-robust and pick up length biases in the preference data (arXiv:2310.03716). Kirk et al. ("Understanding the Effects of RLHF on LLM Generalisation and Diversity", arXiv v1 2023-10-10) compared SFT, reward modeling and RLHF on two base models, summarization and instruction following, and found that RLHF generalizes better than SFT to new inputs, especially under large distribution shift, but significantly reduces output diversity, a quantified form of the "mode collapse" listed as a tractable problem by Casper et al. (arXiv:2310.06452; B05-38). Together they sharpen the over-optimization story (B05-18). A proxy reward can be hacked by verbosity, and the tuned model may be more reliable out of distribution yet less varied, which feeds the "mask" debate (B05-25) and the preference-bias analyses that follow (B05-39). Sources: Singhal et al. · Kirk et al.