GPT-4's report on RLHF, rule-based reward models, and RLHF's effects on exam scores, truthfulness and calibration

GPT-4 (announced 2023-03-14, report arXiv 2023-03-15) was post-trained with RLHF after a pretraining run that, per OpenAI's system card, finished in August 2022 (GPT-4 report).

Date
14 March 2023
Who
OpenAI
Confidence
High (OpenAI's own technical report)
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · Confidence: High (OpenAI's own technical report) GPT-4 (announced 2023-03-14, report arXiv 2023-03-15) was post-trained with RLHF after a pretraining run that, per OpenAI's system card, finished in August 2022 (GPT-4 report). Four facts from the report are the most useful for interviews. (1) RLHF did not change exam capability: averaged across the exams tested, the base model scored 73.7% and the RLHF model 74.0% on multiple-choice portions. (2) RLHF helped truthfulness a lot: on TruthfulQA the base GPT-4 was only slightly better than GPT-3.5, but after RLHF there were large improvements. (3) RLHF hurt calibration: the pretrained model's confidence matched its accuracy well, and post-training reduced this. (4) Rule-based reward models (RBRMs): zero-shot GPT-4 classifiers given a rubric supply an extra reward signal during RLHF to push toward correct refusals and away from over-refusal, an early example of model-assisted feedback alongside human feedback; the report notes that undesired behaviors can arise when labeler instructions were underspecified. The credits list "foundational RLHF and InstructGPT work" as a distinct contribution. Why it matters: it is OpenAI's own statement that RLHF mostly shapes behavior, not core knowledge, supporting the reading that RLHF brings out what the base model already has and teaches little that is new (B05-13, B05-33) while also documenting the calibration trade-off that Schulman's talk argues pretraining provides (B05-32). Sources: OpenAI, GPT-4 Technical Report

Read it in the deep dive