Anthropic's Constitutional AI replaces human harm labels with a written list of principles

Anthropic trained a harmless-but-non-evasive assistant using no human labels for harmful outputs, only a short list of principles.

Date
15 December 2022
Who
Anthropic
People
Yuntao Bai and 50 co-authors (including Jared Kaplan, Dario Amodei, Sam McCandlish, Tom Brown)
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Landmark · Significance: 4/5 · Org(s): Anthropic · People: Yuntao Bai and 50 co-authors (including Jared Kaplan, Dario Amodei, Sam McCandlish, Tom Brown) · Confidence: High Primary sources: arXiv:2212.08073 (v1 2022-12-15). Depth, replications and the later RLAIF literature are in B06; this entry covers only its place in the RLHF lineage.

One-liner. Anthropic trained a harmless-but-non-evasive assistant using no human labels for harmful outputs, only a short list of principles. The model critiques and revises its own answers, and an AI model supplies the preference labels for RL.

Why it happened. The April 2022 HH-RLHF work needed tens of thousands of human red-team comparisons and showed that helpfulness and harmlessness trade off, since helpfulness-trained models were easier to red-team and harmlessness-trained ones tended to become evasive (B05-14). Human labeling was slow and exposed workers to disturbing content (B05-26). Anthropic's framing in the abstract is to enlist AI systems to help supervise other AIs (abstract), the same motive as B05-04.

The idea. Two phases. The supervised phase samples from an initial helpful model on red-team prompts, has it critique and revise its own response against a randomly chosen principle, and fine-tunes on the revisions. The RL phase samples pairs of responses from that model, asks a model which is better according to a principle (optionally with chain-of-thought), trains a preference model on these AI preferences, and runs RL against it. Anthropic named the second step RL from AI Feedback (RLAIF). Human helpfulness labels were still used. The setup used 16 principles, which the authors say were written ad hoc for research and not carefully designed, one sampled at random at each revision step; 182,831 red-team prompts (42,496 human-written plus 140,335 model-generated), with 4 critique-revision pairs sampled per prompt; 135,296 human-written helpfulness prompts; and optional chain-of-thought labeling (paper §3).

Results. On Anthropic's 52B setup, the RL-CAI model was reported as more harmless than the HH-RLHF model and as giving fewer evasive answers, an improvement along the helpfulness/harmlessness frontier, according to crowdworker comparisons and a preference-model evaluation (paper). The intermediate SL-CAI model was rated below both RL models on helpfulness and above the helpful-only RLHF model on harmlessness; the authors also changed the crowdworker instruction to prefer thoughtfully harmless over evasively harmless answers, which affects how the HH-RLHF baseline scores. Company-measured.

How it spread.

Lab/projectResponseRelationshipDateLag vs. original
AnthropicClaude launched; Anthropic says CAI, with around ten principles it had not yet published, is what sets the model apart (company-claimed; TechCrunch, 2023-03-14)originator's own deployment2023-03-1413 weeks
AnthropicPublishes the principle list as "Claude's constitution" (B05-32c)originator's follow-up2023-05-0921 weeks
Sun et al. (academic and industry authors)Principle-Driven Self-Alignment ("Dromedary") on LLaMA-65B with 16 generic principles, under 300 lines of human annotation, SFT only (arXiv:2305.03047)parallel, principle-driven, no RL2023-05-0420 weeks
OpenAIGPT-4's rule-based reward models (B05-29)parallel; no evidence it derives from CAI2023-03-1413 weeks
MetaLlama 2 paper cites CAI and mentions "RL from AI feedback" as the idea of using a model to rank outputs (Llama 2)cites it; did not adopt it2023-07-1831 weeks
GoogleRLAIF vs. RLHF, the first independent lab-scale test that AI preference labels can match human ones (B05-38a)independent replication/extension2023-09-0137 weeks
AnthropicCollective Constitutional AI, with about 1,000 Americans' input to the principles (B05-32c)originator's follow-up2023-10-1744 weeks
AnthropicA new, much longer constitution (B05-42f)originator's successor2026-01-21about 162 weeks

The paper and Anthropic's own product carried it; closer followers are in B06. Claude's launch is the originator's own deployment and GPT-4's rule-based reward models are parallel work, so neither counts as a follower.

Why it mattered. It was the first widely cited route to reduce the human-labor share of alignment data, with written principles as the human input, and it moved the conversation from "who labels" to "who writes the constitution".

Nuance, controversy and myths. CAI does not remove humans, since people write the principles, supply helpfulness preference labels, and red-team. "AI feedback" does not mean the model chooses its own values. Whether AI-generated preference labels carry the labeler biases of the generating model is an open question.

Interview kit.

  • 30-second version: Constitutional AI replaces human harm labels with principles. The model revises its own outputs against rules, and then an AI judge produces the preferences used for RL (RLAIF).
  • Likely follow-ups: Is that RLHF? → It's RLHF's cousin, with the same pipeline except that the preference labels for harmlessness come from a model. Is it cheaper? → It shifts cost from labelers to principle design and compute.
  • Common mistake: Saying Claude was trained without any human feedback.
  • Connect it to: B05-14, B05-30, B06, B24.

Sources. 1. Bai et al., arXiv:2212.08073 · 2. Llama 2 paper · 3. TechCrunch, 2023-03-14 · 4. Sun et al., arXiv:2305.03047

Read it in the deep dive