Measuring Progress on Scalable Oversight for Large Language Models
Anthropic ran its first empirical test of supervising systems that may outperform us, the idea behind B05-04.
- Date
- 4 November 2022
- Who
- Anthropic
- People
- Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez and 43 co-authors, including Amanda Askell, Dario Amodei
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 2/5 · Org(s): Anthropic · People: Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez and 43 co-authors, including Amanda Askell, Dario Amodei · Confidence: High Anthropic ran its first empirical test of supervising systems that may outperform us, the idea behind B05-04. The paper proposes studying it on tasks where human specialists succeed but unaided humans and current AI fail, then runs a proof of concept on two question-answering tasks (MMLU and time-limited QuALITY). People who chatted with an unreliable LLM assistant, a deliberately trivial baseline, did substantially better than both the model alone and their own unaided performance (arXiv:2211.03540, v1 2022-11-04). It is a modest result, mainly a method for studying the problem with present models, and it complements OpenAI's critique work (B05-14a) and the later weak-to-strong results (B05-40a). Sources: Bowman et al.