Anthropic publishes HH-RLHF and red-teaming results, with iterated online RLHF at 52B
Anthropic's first RLHF paper trained 52B helpful-and-harmless assistants and showed that updating the preference model and policy weekly with fresh human data improved both.
- Date
- 12 April 2022
- Who
- Anthropic
- People
- Yuntao Bai and Jared Kaplan (corresponding) and 29 other authors, including Amanda Askell, Dario Amodei, Tom Brown, Jack Clark, Chris Olah
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Landmark · Significance: 4/5 · Org(s): Anthropic · People: Yuntao Bai and Jared Kaplan (corresponding) and 29 other authors, including Amanda Askell, Dario Amodei, Tom Brown, Jack Clark, Chris Olah · Confidence: High Primary sources: arXiv:2204.05862 (v1 2022-04-12) · red teaming, arXiv:2209.07858 (v1 2022-08-23)
One-liner. Anthropic's first RLHF paper trained 52B helpful-and-harmless assistants and showed that updating the preference model and policy weekly with fresh human data improved both.
Why it happened. After HHH (B05-10) the open question was whether the recipe holds up with a harmlessness objective that conflicts with helpfulness and with a different, more adversarial kind of labeler task, red teaming (crowdworkers trying to make the model say harmful things). Inference: the company was about a year old and (as far as I can find) had no public API of its own until March 2023, unlike OpenAI with its customers' prompts, so it collected its own conversations with crowdworkers through an in-house interface, using its large pretrained models.
The idea. Train preference models (PMs) on human comparisons for helpfulness and for harmlessness (red-teaming), then RL-fine-tune with the PM score as reward; in an iterated online mode, retrain the PM and policy about weekly and send the new policy back to crowdworkers.
How it works. The initial helpfulness data was collected with a context-distilled LM; the core base dataset has 44k helpfulness and 42k red-teaming comparisons; a rejection-sampling dataset adds 52k helpfulness and 2k red-team comparisons; the online dataset adds 22k helpfulness comparisons across about five weeks (appendix); RL prompts were 137k human-written plus 369k model-generated. RL reward is roughly linear in the square root of the KL divergence from the initial policy. Labor. About 30 "select" workers, roughly half hired through Upwork and half from US MTurk workers with a Masters qualification; MTurk workers accounted for 80 to 85% of the select group's comparison data; a Slack channel supported daily communication; the authors say workers were paid well above California's minimum wage but note MTurk pay is per task (paper App. D). The companion red-teaming study used 324 US-based crowdworkers (307 MTurk, 17 Upwork), paid $7.50 to $9.50 per set of five conversations; the authors report this is at or above California minimum wage (red-teaming paper).
Results. Alignment training improved performance on most NLP evaluations for 13B and 52B models (an "alignment bonus") and was compatible with coding and summarization skills, while smaller models paid a real "alignment tax" (paper). Online iterated training improved crowdworker-judged quality. Helpfulness and harmlessness trade off, but the trade-off shrinks with PM size. In the red-teaming paper, across 2.7B, 13B and 52B parameters and four model types, RLHF models became harder to red-team with scale while the others showed flat trends, and 38,961 attacks were released as a dataset (abstract). The helpful-and-harmless comparisons were released publicly (Hugging Face dataset Anthropic/hh-rlhf) and became a standard open preference set.
How it spread. Anthropic published this work but did not ship a product. Its first assistant, Claude, was built and held back during 2022 (B05-30). The released HH data was used as an open preference dataset, including in Llama 2's reward-model training (which also used OpenAI's summarization and WebGPT comparisons) (Llama 2 Table 6).
Why it mattered. It was the second independent, detailed account of RLHF at scale, from a lab with a different safety philosophy, and the one that released data. It also introduced iterated online RLHF as a practice.
Nuance, controversy and myths. "Helpful and harmless" are treated as separate rewards that fight; the later Constitutional AI paper tries to reduce human labeling of harmlessness (B05-21). The paper's labor model (a few dozen high-priority workers on Slack, plus a large pool paid per task) contrasts with the Kenya and Remotasks reporting in B05-26.
Interview kit.
- 30-second version: Anthropic's April 2022 paper applied RLHF to 52B models with separate helpfulness and harmlessness preference models, updating them weekly with new human data, and released the data.
- Likely follow-ups: Does safety training make models worse? → For small models yes, for 13B and 52B it was roughly free or beneficial in this paper. Where is the data from? → About 30 select workers (Upwork and MTurk) plus a larger pool.
- Common mistake: Saying Anthropic invented RLHF; it independently implemented and published it (B05-02).
- Connect it to: B05-10, B05-13, B05-21, B22.
Sources. 1. Bai et al., arXiv:2204.05862 · 2. Ganguli et al., arXiv:2209.07858 · 3. Touvron et al., arXiv:2307.09288