OpenAI trains self-critiquing models to help humans evaluate summaries

This was OpenAI's first empirical test, on language models, of the scalable-oversight idea in B05-04 that a model can help people judge outputs they could not easily judge alone.

Date
12 June 2022
Who
OpenAI
People
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, Jan Leike
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · People: William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, Jan Leike · Confidence: High This was OpenAI's first empirical test, on language models, of the scalable-oversight idea in B05-04 that a model can help people judge outputs they could not easily judge alone. The team fine-tuned models by behavior cloning to write natural-language critiques of topic-based summaries. Critiques helped human raters find flaws they would otherwise have missed, including naturally occurring flaws in model and human summaries and deliberately planted misleading ones; larger models wrote more helpful critiques, and larger models could use their own critiques to refine their summaries. The authors also found a gap between what models can generate, discriminate and critique, which suggests that even large models may hold knowledge they do not articulate as critiques. They released their datasets (arXiv:2206.05802, v1 2022-06-12). Later work connects to it. Anthropic's critique-and-revise step in Constitutional AI is a related idea (Inference; B05-21), Anthropic's own proof of concept for oversight followed in November (B05-18a), and OpenAI's later CriticGPT is the code-review descendant (B05-40c). Sources: Saunders et al.

Read it in the deep dive