OpenAI reports clinician-reviewed changes to ChatGPT's responses in sensitive conversations

OpenAI reported that it worked with more than 170 mental-health experts, drawn from a network of nearly 300 physicians and psychologists, to improve how ChatGPT handles psychosis or mania, self-harm…

Date
27 October 2025
Who
OpenAI
Confidence
Medium (company-reported figures, partly graded by clinicians OpenAI recruited;
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · Confidence: Medium (company-reported figures, partly graded by clinicians OpenAI recruited; no independent audit) OpenAI reported that it worked with more than 170 mental-health experts, drawn from a network of nearly 300 physicians and psychologists, to improve how ChatGPT handles psychosis or mania, self-harm and suicide, and emotional reliance on AI, and that the updated default model returned responses falling short of its desired behavior 65% to 80% less often across those domains. Against GPT-4o, expert graders found 39% fewer undesired answers on challenging mental-health conversations (n=677) and 42% fewer on emotional-reliance conversations (n=507). The post also gave first prevalence estimates, which it described as hard to measure. About 0.07% of weekly active users and 0.01% of messages showed possible signs of psychosis or mania, about 0.15% of weekly users had explicit indicators of suicidal planning or intent, and about 0.15% showed heightened emotional attachment to ChatGPT (OpenAI, 2025-10-27, read from an archived copy). It matters here as the first lab-scale quantification I found of how often the sycophancy-and-dependence problem shows up in production traffic. It is also a post-training fix built on clinician-written ideal responses and clinician grading, a labor model that sits between the crowd pools of B05-26 and the expert vendors of B05-42. Related entries are B05-41 and B05-42b. Sources: OpenAI

Read it in the deep dive