Anthropic publishes A General Language Assistant as a Laboratory for Alignment

Anthropic's first major paper defined the "helpful, honest, harmless" (HHH) target and showed that ranked preference modeling beats imitation learning and often scales more favorably.

Date
1 December 2021
Who
Anthropic
People
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Ben Mann (core research); 22 authors including Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Jared Kaplan
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Landmark · Significance: 4/5 · Org(s): Anthropic · People: Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Ben Mann (core research); 22 authors including Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Jared Kaplan · Confidence: High Primary sources: arXiv:2112.00861 (v1 2021-12-01)

One-liner. Anthropic's first major paper defined the "helpful, honest, harmless" (HHH) target and showed that ranked preference modeling beats imitation learning and often scales more favorably. The paper set up Anthropic's RLHF work.

Why it happened. The authors included people who had come from OpenAI, namely Amodei and Brown (co-authors of the 2017 and 2019 papers above, B05-02, B05-05) and Askell, whom the InstructGPT paper lists as formerly at OpenAI (InstructGPT footnote). The paper states what the new company would do, which was to study alignment on a general-purpose text assistant with simple baselines. Anthropic's own Series A announcement (2021-05-28, $124 million) already listed GPT-3, "Concrete Problems in AI Safety" and "Learning from Human Preferences" among the team's earlier research and said it would build ways to integrate human feedback more tightly into AI systems (Anthropic).

The idea. Use "helpful, honest, harmless" as a simple, memorable definition of an aligned assistant; study cheap interventions (prompting, "context distillation" which bakes the prompt into the weights) and then study scaling trends for training objectives.

How it works. The paper tested models from 10M to 52B parameters. Its findings are that (1) a long HHH prompt improves alignment evaluations, increasingly with size, and imposes little or no "alignment tax" on large models, tested on Codex-style coding evaluations; (2) ranked preference modeling (train a model to score responses so ranking matches human rankings) beats imitation learning and often scales more favorably, while binary discrimination performs and scales like imitation; (3) a preference-model pre-training (PMP) stage on large public ranked data (Stack Exchange, Reddit, reverted Wikipedia vandalism) before fine-tuning on small human-preference sets improves sample efficiency (abstract and Fig. 4). The paper does not train an RL-tuned assistant; it states RLHF as the next step.

Results. Prompted behavior improved on the authors' HHH evaluations and benefits grew with scale; ranked preference models beat imitation learning on the evaluations the authors tested. The paper also frames preference models as the basis for debate- and amplification-style proposals (B05-04). Lab-measured. The paper's own "alignment tax" finding applies to prompting only.

How it spread. InstructGPT (2022) cites it as concurrent work and adopts HHH as its target (helpful, honest, harmless) (InstructGPT §1-2). Anthropic's April 2022 paper (B05-14) is the RL follow-up; HHH became the common vocabulary, later reworked into constitution and spec approaches (B06).

Why it mattered. It fixed the target and the vocabulary of HHH, preference models, context distillation and alignment tax. Inference: it also shows that Anthropic's alignment recipe existed within months of the company's founding because the people carried the methods (B05-02).

Nuance, controversy and myths. The acronym "RLHF" appears only a few times (as future work, citing Christiano et al. 2017); it was not yet a household word. "Alignment tax" circulated in the alignment community before this paper (Christiano is commonly cited for it; earlier origin not traced), and the paper's no-tax claim is specific to prompting at large scale; the later HH paper found small models paid a real tax (B05-14).

Interview kit.

  • 30-second version: Anthropic's December 2021 paper defined helpful-honest-harmless, showed ranked preference models beat imitation learning, and introduced context distillation and preference-model pretraining.
  • Likely follow-ups: Where does "HHH" come from? → This paper; InstructGPT then used the same trio. What is an alignment tax? → Capability lost by aligning; for large models prompting costs little.
  • Common mistake: Saying Anthropic's first paper trained Claude with RLHF; it studied prompting and preference modeling.
  • Connect it to: B05-13, B05-14, B05-21.

Sources. 1. Askell et al., arXiv:2112.00861 · 2. Ouyang et al.

Read it in the deep dive