Hugging Face's StackLLaMA, an open RLHF recipe on LLaMA-7B
A hands-on guide from Hugging Face ran the three InstructGPT steps on Meta's LLaMA-7B using Stack Exchange data.
- Date
- 5 April 2023
- Who
- Hugging Face
- People
- Edward Beeching, Kashif Rasul, Younes Belkada, Lewis Tunstall, Leandro von Werra, Nazneen Rajani, Nathan Lambert
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): Hugging Face · People: Edward Beeching, Kashif Rasul, Younes Belkada, Lewis Tunstall, Leandro von Werra, Nazneen Rajani, Nathan Lambert · Confidence: High A hands-on guide from Hugging Face ran the three InstructGPT steps on Meta's LLaMA-7B using Stack Exchange data. The steps were supervised fine-tuning on questions and answers, a reward model trained to predict which of two answers was preferred, and PPO through Hugging Face's TRL library, with 8-bit weights and LoRA adapters through PEFT to fit memory and 3 x 8 A100-80GB GPUs for the RL phase (HF blog, 2023-04-05). It matters because it was an early public, end-to-end, reproducible example of RLHF outside a frontier lab, and because it documented the practical failure modes. The policy learned that emitting code blocks (common on Stack Exchange) raised the reward, an instance of reward exploitation (B05-18), and the KL penalty, which should never be negative, went negative because of forced token generation during batched inference. It came three months before Llama 2 documented the industrial version (B05-37) and is part of the open-tooling story in B05-17. Sources: Hugging Face, StackLLaMA · TRL repo