OpenAI's post "How should AI systems behave, and who should decide?"

Weeks after ChatGPT, OpenAI published an explicit public description of how the fine-tuning stage shapes behavior.

Date
16 February 2023
Who
OpenAI
Confidence
High (the post's contents); Low (the surrounding controversy, whose coverage I d
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · Confidence: High (the post's contents); Low (the surrounding controversy, whose coverage I did not open) Weeks after ChatGPT, OpenAI published an explicit public description of how the fine-tuning stage shapes behavior. Pretraining on large text gives models grammar, facts and some reasoning along with biases in the text, and fine-tuning then uses human reviewers who follow guidelines OpenAI provides, rating a range of model outputs without per-input instructions. The post shares the part of those guidelines on political and controversial topics, says reviewers should not favor any political group, treats biases that emerge anyway as bugs and not as intended features, promises clearer reviewer instructions and aggregated reviewer demographics, and lays out three building blocks, which are better default behavior, user customization within broad bounds (warning that unlimited customization risks "sycophantic AIs that mindlessly amplify people's existing beliefs"), and public input on defaults and hard bounds. It also names rule-based rewards and Constitutional AI as methods it was building on to make fine-tuning more understandable and controllable (OpenAI, 2023-02-16, read from an archived copy). It belongs here because, on my reading (Inference), it is the lab-side answer to worries that RLHF encodes the reviewers' politics; it names "sycophantic" AI as a risk two months after Anthropic first measured sycophancy (B05-22); and it ties reviewer demographics to the labor entries (B05-26). I did not open the early-2023 press coverage of the political-bias claims, so that dispute is not characterized here (Backlog). Sources: OpenAI

Read it in the deep dive