Kalai et al. argue that hallucination comes from pretraining pressure and exam-style benchmark grading

The OpenAI and Georgia Tech paper "Why language models hallucinate" argues that hallucinations originate as errors in binary classification of whether an output is valid, so pretraining produces them…

Date
4 September 2025
Who
OpenAI, Georgia Tech
People
Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang
Confidence
High (the paper's argument); Medium (how much it explains deployed behavior, whi
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI, Georgia Tech · People: Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang · Confidence: High (the paper's argument); Medium (how much it explains deployed behavior, which the paper does not test at scale) The OpenAI and Georgia Tech paper "Why language models hallucinate" argues that hallucinations originate as errors in binary classification of whether an output is valid, so pretraining produces them through statistical pressure. It also argues that post-training fails to remove them because most benchmarks grade like exams, rewarding a lucky guess over an honest "I don't know". The proposed fix is socio-technical, changing how existing benchmarks score uncertainty, and the paper does not propose new hallucination tests (arXiv:2509.04664, v1 2025-09-04). It is the 2025 sequel to Schulman's 2023 argument about calibration and RL (B05-32). Both locate the problem in rewards that do not penalize confident errors more than abstentions, and it echoes GPT-4's observation that post-training reduced calibration (B05-29). Inference. The same incentive logic applies to preference models that reward agreeable answers (B05-39). Sources: Kalai et al.

Read it in the deep dive