"How is ChatGPT's behavior changing over time?" by Chen, Zaharia and Zou, the paper behind the "GPT-4 got dumber" story

Chen, Zaharia and Zou compared the March 2023 and June 2023 versions of GPT-3.5 and GPT-4. Version warning.

Date
18 July 2023
Who
Stanford, UC Berkeley
People
Lingjiao Chen, Matei Zaharia, James Zou
Confidence
High (the paper's measurements); Low (any causal link to RLHF, which the paper d
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): Stanford, UC Berkeley · People: Lingjiao Chen, Matei Zaharia, James Zou · Confidence: High (the paper's measurements); Low (any causal link to RLHF, which the paper does not make) Chen, Zaharia and Zou compared the March 2023 and June 2023 versions of GPT-3.5 and GPT-4. Version warning. The first arXiv version (2023-07-18) covered four tasks and reported GPT-4's accuracy at identifying prime numbers falling from 97.6% to 2.4%; the current abstract (v3, 2023-10-31) covers seven tasks and reports 84% falling to 51% on prime-versus-composite questions, partly explained by the June model following chain-of-thought prompts less. Both versions report that GPT-3.5 improved on the same task, that GPT-4 became less willing to answer sensitive questions, and that both models made more formatting mistakes in code generation; the revised version adds that GPT-4 did better on multi-hop questions in June. The authors read the pattern as evidence that GPT-4's ability to follow user instructions had decreased and called for continuous monitoring of hosted models (arXiv:2307.09009). Arvind Narayanan and Sayash Kapoor replied the next day that the paper shows behavior drift, not capability loss. In their view pretraining capabilities persist while fine-tuning changes how they are expressed, and the test used only prime numbers; when they added 500 composites, all four model versions did equally poorly, which they read as guessing from calibration patterns (AI Snake Oil / Normal Tech, 2023-07-19). It belongs here because it is the origin of the "post-training makes models worse" narrative and a standard interview trap. The paper observes changes in a hosted service whose updates are opaque and does not isolate RLHF, so it gives no evidence about an alignment tax in the sense of B05-13 or B05-29. Sources: Chen et al. · Narayanan and Kapoor

Read it in the deep dive