AI2, CarperAI, Hugging Face and Microsoft release open RLHF tooling for large models
Open RLHF code long predates 2022. OpenAI's own Ziegler et al. code (openai/lm-human-preferences) was created on 2019-09-14 and Hugging Face's TRL library (PPO for Transformers models) on 2020-03-27,…
- Date
- 3 October 2022
- Who
- AI2/Fraunhofer, CarperAI, Hugging Face, LAION, Microsoft
- Confidence
- High (repository and paper dates); Medium (DeepSpeed-Chat date and claims from a
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): AI2/Fraunhofer, CarperAI, Hugging Face, LAION, Microsoft · Confidence: High (repository and paper dates); Medium (DeepSpeed-Chat date and claims from a secondary report) Open RLHF code long predates 2022. OpenAI's own Ziegler et al. code (openai/lm-human-preferences) was created on 2019-09-14 and Hugging Face's TRL library (PPO for Transformers models) on 2020-03-27, per GitHub repository metadata (lm-human-preferences; TRL). What arrived between InstructGPT and Llama 2 was tooling for large models and how-to write-ups that made RLHF reproducible.
RL4LMs and the GRUE benchmark (AI2, Fraunhofer IAIS, UW; arXiv 2022-10-03; repository created 2022-08-18) offered an open library for RL on HuggingFace models, a benchmark of six reward-supervised generation tasks, and NLPO, which the authors report was more stable than PPO; they concluded RL techniques generally align LMs to human preferences better than supervised methods (Ramamurthy et al.). trlX (CarperAI) was created on GitHub on 2022-10-03 (trlX); InfoQ reported in January 2023 that LAION (OpenAssistant), CarperAI (trlX) and Phil Wang (an independent developer) had released open implementations (InfoQ, 2023-01).
The first widely reproduced open recipe on a LLaMA model was Hugging Face's StackLLaMA (B05-30b); AlpacaFarm added a cheap simulator for RLHF research (B05-33a). Hugging Face's "Illustrating Reinforcement Learning from Human Feedback (RLHF)" explainer was published 2022-12-09 (blog), nine days after ChatGPT. Microsoft's DeepSpeed-Chat (2023-04-12) packaged the three InstructGPT steps into one pipeline and claimed 15x throughput over earlier systems (Gigazine report). The gating factors were PPO's engineering cost and human data (B05-03). Fully open RLHF did not reach production quality until Llama 2 documented it (B05-37). Sources: RL4LMs paper · HF blog · InfoQ · DeepSpeed-Chat coverage · lm-human-preferences repo · TRL repo · trlX repo