OpenAI trains CriticGPT, an LLM critic that catches bugs in model-written code

OpenAI's paper opens by stating that RLHF is fundamentally limited by humans' capacity to evaluate model output, and trains "critic" models, themselves trained with RLHF, to write natural-language…

Date
28 June 2024
Who
OpenAI
People
Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trębacz, Jan Leike
Confidence
High (the paper's results; company-measured)
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · People: Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trębacz, Jan Leike · Confidence: High (the paper's results; company-measured) OpenAI's paper opens by stating that RLHF is fundamentally limited by humans' capacity to evaluate model output, and trains "critic" models, themselves trained with RLHF, to write natural-language feedback on problems in code from real assistant tasks. On code with naturally occurring LLM errors, the critiques were preferred over human critiques in 63% of cases, human evaluation found the models caught more bugs than contractors paid for code review, and the critics found hundreds of errors in ChatGPT training data that had been rated flawless, even though most of those tasks were not code; limitations include hallucinated bugs, and human-and-critic teams caught a similar number of bugs to critics alone while hallucinating less (arXiv:2407.00215, v1 2024-06-28). It is a direct admission that human RLHF labels are noisy at the high end, and the code-review descendant of the critique idea in B05-14a and of the oversight agenda in B05-04. Last author Jan Leike joined Anthropic in May 2024 (CNBC), so the paper reads as the tail of Leike's OpenAI line of work (Inference). Sources: McAleese et al.

Read it in the deep dive