OpenAI trains self-critiquing models to help humans evaluate summaries
This was OpenAI's first empirical test, on language models, of the scalable-oversight idea in B05-04 that a model can help people judge outputs they could not easily judge alone.
- Date
- 12 June 2022
- Who
- OpenAI
- People
- William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, Jan Leike
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · People: William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, Jan Leike · Confidence: High This was OpenAI's first empirical test, on language models, of the scalable-oversight idea in B05-04 that a model can help people judge outputs they could not easily judge alone. The team fine-tuned models by behavior cloning to write natural-language critiques of topic-based summaries. Critiques helped human raters find flaws they would otherwise have missed, including naturally occurring flaws in model and human summaries and deliberately planted misleading ones; larger models wrote more helpful critiques, and larger models could use their own critiques to refine their summaries. The authors also found a gap between what models can generate, discriminate and critique, which suggests that even large models may hold knowledge they do not articulate as critiques. They released their datasets (arXiv:2206.05802, v1 2022-06-12). Later work connects to it. Anthropic's critique-and-revise step in Constitutional AI is a related idea (Inference; B05-21), Anthropic's own proof of concept for oversight followed in November (B05-18a), and OpenAI's later CriticGPT is the code-review descendant (B05-40c). Sources: Saunders et al.