Meta releases Llama 2-Chat, the first open-weights chat model with a detailed RLHF recipe

Meta released chat models from 7B to 70B with weights usable commercially, and the paper laid out an SFT-plus-RLHF pipeline (1.4M human comparisons, two reward models, five rounds of rejection…

Date
18 July 2023
Who
Meta (GenAI)
People
Hugo Touvron and Thomas Scialom (corresponding authors); Louis Martin, Kevin Stone, Guillem Cucurull, Ruan Silva, Naman Goyal (science leadership); Sergey Edunov, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Landmark · Significance: 5/5 · Org(s): Meta (GenAI) · People: Hugo Touvron and Thomas Scialom (corresponding authors); Louis Martin, Kevin Stone, Guillem Cucurull, Ruan Silva, Naman Goyal (science leadership); Sergey Edunov, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic · Confidence: High Primary sources: arXiv:2307.09288 (v1 2023-07-18) · Meta announcement, 2023-07-18

One-liner. Meta released chat models from 7B to 70B with weights usable commercially, and the paper laid out an SFT-plus-RLHF pipeline (1.4M human comparisons, two reward models, five rounds of rejection sampling with PPO added in the last) in enough detail for others to copy. It was the first open-weights chat model with such a recipe. Detailed public RLHF accounts already existed (B05-13, B05-14).

Why it happened. Meta's LLaMA 1 (2023-02-24) was a base model with no assistant tuning; its leak and the Alpaca and Vicuna fine-tunes of it (B05-28) showed demand and exposed the gap in tuning quality. Meta's earlier public assistant attempts had drawn criticism (B05-15); the team that wrote Llama 2 was largely a Meta GenAI group that overlapped with Galactica's authors. The Llama 2 post says there had been more than 100,000 requests for access to Llama 1 (Meta).

The idea. Do ChatGPT-style post-training in the open and document the numbers, so the open community can reproduce it and enterprises can adopt it.

How it works. Pretraining: 2 trillion tokens, 4k context. SFT: start from public instruction data (Flan), then collect only about 27,540 high-quality vendor-written examples; the authors say that setting aside millions of third-party examples in favor of fewer, better ones improved results considerably, that different vendors yield markedly different downstream quality, and that SFT-model samples were often competitive with human-written ones, so they shifted annotation effort to preferences. Reward models: a binary-comparison protocol (annotators write a prompt, pick the better of two responses from different model variants and rate the strength of preference); over 1 million comparisons, with Meta's own set at 1,418,091 comparisons plus open sets (Anthropic Helpful 122,387; OpenAI Summarize 176,625; WebGPT, StackExchange and others); two reward models, helpfulness and safety, with a margin term in the ranking loss tied to preference strength. Optimization: five successive versions (RLHF-V1 to V5) as new preference batches arrived. Through V4 the team used only rejection sampling fine-tuning (best of K samples from the largest 70B model, with smaller models trained on its outputs); only for the final version, V5, did they apply PPO on top of the rejection-sampling checkpoint (the paper's Figure 11 compares "RLHF-v5 (with PPO)" with "RLHF-v5 (no PPO)"); plus Ghost Attention for multi-turn system-instruction consistency. Red teaming involved over 350 people including contractors and external vendors (paper §3, §4). The paper also reports that the 34B model's release was delayed for lack of time to red-team.

Results. In a human evaluation on about 4,000 helpfulness prompts with three raters each, Llama 2-Chat 70B had a 36% win rate and 31.5% tie rate against ChatGPT, and beat comparably sized open chat models by wide margins (Llama 2-Chat 34B above 75% against Vicuna-33B and Falcon-40B). Meta's own reward models showed the chat model surpassing ChatGPT on both axes by RLHF-V3, and GPT-4-as-judge gave Llama 2-Chat a win rate above 60% (company-measured; the authors note the reward-model comparison may favor their own model).

How it spread.

Lab/projectResponseRelationshipDateLag vs. Llama 2
AlibabaQwen-7B and Qwen-7B-Chat released; the report (arXiv 2023-09-28) describes a reward model and PPO, but the README says the RLHF chat models were not released (B05-30a)parallel, independent (released 16 days after Llama 2; I did not establish whether the first checkpoint had an RLHF stage)2023-08-03 (repo/README)2 weeks
MistralMistral 7B Instruct was fine-tuned on public instruction datasets with "no tricks, no proprietary data" per Mistral's post; the paper compares it with Llama 2 13B-Chatindependent; a cheap SFT demonstration with no RLHF2023-09-27 (announcement; arXiv 2023-10-10)10 weeks
DeepSeekDeepSeek LLM Chat used about 1.5M SFT instances then DPOindependent; chose DPO over PPO (B05-35)2023-11-29 (repo; arXiv 2024-01-05)19 weeks
GoogleGemini describes SFT, reward modeling and RLHF (B05-40)own pipeline, documented late2023-12-0620 weeks

Open models picked up the recipe quickly because Meta published both the paper and the weights. Real RLHF reproduction was slower because of the human data bill (Meta's 1.4M comparisons) and PPO's cost. Several open follow-ups used DPO instead, Zephyr and DeepSeek LLM Chat among them, and Zephyr also replaced human labels with AI feedback (B05-35, B05-39a).

Why it mattered. It made RLHF a documented, reproducible engineering practice and gave "open weights" its first strong assistant. It also contains a notable open statement that RLHF can lift models past annotator ability. The authors argue that humans are better at comparing than writing, so the reward signal can exceed the writing ability of the annotators, and that supervised data may no longer be the gold standard (§5.1).

Nuance, controversy and myths. (1) The 1.4M figure counts binary comparisons, which is a different unit from conversations or labelers. (2) The labor behind it is unnamed. The paper thanks annotators and internal annotation leads but does not identify vendors, locations or pay (B05-26). (3) "Open source" is a loose label for weights under a custom license with usage conditions (B19). (4) Llama 2-Chat used rejection sampling and PPO; the DPO-era recipes came after.

Interview kit.

  • 30-second version: Llama 2-Chat (July 2023) is Meta's open assistant. It used about 27.5k high-quality SFT examples, over 1M human comparisons, separate helpfulness and safety reward models, and five rounds of iteration with rejection sampling and, in the last round only, PPO on top. It made RLHF a reproducible recipe.
  • Likely follow-ups: "Did it beat ChatGPT?" → In Meta's human eval 70B had 36% wins and 31.5% ties, so it was competitive without being dominant. "What did they say about humans vs models?" → Comparing is easier than writing, so RLHF can surpass annotator writing quality. "Why did people then switch to DPO?" → Cost and stability (B05-35).
  • Common mistake: Saying Llama 2 is "just SFT" or that Meta used a single reward model.
  • Connect it to: B05-13, B05-14, B05-33, B19.

Sources. 1. Touvron et al., arXiv:2307.09288 · 2. Meta, "Llama 2", 2023-07-18 · 3. Mistral 7B, arXiv:2310.06825 · 4. Qwen report · 5. DeepSeek LLM

Read it in the deep dive