Hugging Face's StackLLaMA, an open RLHF recipe on LLaMA-7B

A hands-on guide from Hugging Face ran the three InstructGPT steps on Meta's LLaMA-7B using Stack Exchange data.

Date
5 April 2023
Who
Hugging Face
People
Edward Beeching, Kashif Rasul, Younes Belkada, Lewis Tunstall, Leandro von Werra, Nazneen Rajani, Nathan Lambert
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): Hugging Face · People: Edward Beeching, Kashif Rasul, Younes Belkada, Lewis Tunstall, Leandro von Werra, Nazneen Rajani, Nathan Lambert · Confidence: High A hands-on guide from Hugging Face ran the three InstructGPT steps on Meta's LLaMA-7B using Stack Exchange data. The steps were supervised fine-tuning on questions and answers, a reward model trained to predict which of two answers was preferred, and PPO through Hugging Face's TRL library, with 8-bit weights and LoRA adapters through PEFT to fit memory and 3 x 8 A100-80GB GPUs for the RL phase (HF blog, 2023-04-05). It matters because it was an early public, end-to-end, reproducible example of RLHF outside a frontier lab, and because it documented the practical failure modes. The policy learned that emitting code blocks (common on Stack Exchange) raised the reward, an instance of reward exploitation (B05-18), and the KL penalty, which should never be negative, went negative because of forced token generation during batched inference. It came three months before Llama 2 documented the industrial version (B05-37) and is part of the open-tooling story in B05-17. Sources: Hugging Face, StackLLaMA · TRL repo

Read it in the deep dive