John Schulman's Berkeley talk arguing that supervised fine-tuning teaches hallucination and RL may not

In a UC Berkeley EECS colloquium, "Reinforcement Learning from Human Feedback: Progress and Challenges", Schulman made the clearest public argument for why the ChatGPT pipeline uses RL as well as…

Date
19 April 2023
Who
OpenAI
People
John Schulman (introduced by Pieter Abbeel as chief architect of ChatGPT)
Confidence
High (first-hand talk; the talk was on 2023-04-19, an EECS Colloquium at Sutardj
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · People: John Schulman (introduced by Pieter Abbeel as chief architect of ChatGPT) · Confidence: High (first-hand talk; the talk was on 2023-04-19, an EECS Colloquium at Sutardja Dai Hall, 5-6 p.m.; the Berkeley News transcript was published 2023-04-24) In a UC Berkeley EECS colloquium, "Reinforcement Learning from Human Feedback: Progress and Challenges", Schulman made the clearest public argument for why the ChatGPT pipeline uses RL as well as supervised learning (event page). Behavior cloning (Schulman's term for supervised fine-tuning) teaches hallucination because the correct target depends on what the model knows, which the labeler cannot see. Training on correct answers the model lacks teaches it to guess, and training it to say "I don't know" teaches it to withhold what it has. Schulman predicted that models fine-tuned on ChatGPT outputs would hallucinate more than the original, about five weeks before Berkeley researchers reported that imitation closes little of the capability gap (B05-34). His claim was that pretrained models are calibrated because they minimize log loss, so RL with a reward for correct answers, a penalty for wrong ones and a neutral value for refusing can teach thresholded answering; in the talk's experiment RL against an oracle reward learned this and RL against a learned reward model "basically worked" but did worse. He gave these caveats. Ranking-based reward models say only how confident they are that one answer beats another and say nothing about how bad a factual error is; detailed labeling interfaces for factuality still yielded a single bit of preference; labelers cannot catch every error in long answers; and RLHF optimizes for what sounds convincing, which can differ from what is true (Berkeley transcript). This frames the open problem that later verifiable-reward RL addressed (B07, B08); Kalai et al.'s 2025 argument about guessing versus abstaining is its sequel (B05-42c). Sources: Berkeley Talks transcript, 2023-04-24 · EECS colloquium page · recording

Read it in the deep dive