text-davinci-001/002/003 and the GPT-3.5 series, and which of them used RLHF
OpenAI's API model names did not say which models were trained with RLHF. In the lineage described by OpenAI's model index (the page itself is no longer retrievable; I relied on a Fudan University…
- Date
- 28 November 2022
- Who
- OpenAI
- Confidence
- High (for the method labels); Medium (exact release days)
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · Confidence: High (for the method labels); Medium (exact release days) OpenAI's API model names did not say which models were trained with RLHF. In the lineage described by OpenAI's model index (the page itself is no longer retrievable; I relied on a Fudan University analysis that reproduces it and on a contemporaneous LessWrong thread), the order is base davinci (2020); text-davinci-001, an InstructGPT model trained with "FeedME", meaning supervised fine-tuning on human demonstrations plus model samples that labelers rated highly (Jan Leike told the LessWrong poster the sample-generating models were a mix of earlier ones); code-davinci-002, the code model that is the base of the GPT-3.5 series; text-davinci-002, FeedME on that base; text-davinci-003, trained with PPO against a reward model trained from human comparisons; then gpt-3.5-turbo, the chat-optimized model (Ye et al., arXiv:2303.10420; LessWrong/Janus update, 2022-11-19, with 2022-11-30 comments). The InstructGPT blog had already noted that the deployed API models used a similar but slightly different method than the paper's (blog footnote). text-davinci-003 was announced on Monday 2022-11-28; Jan Leike described it as largely equivalent to the InstructGPT models but not identical, scoring higher on human preference without being more capable (The Decoder, 2022-11-29; VentureBeat). Why it matters for interviews: "text-davinci-002 is RLHF" was a widespread community assumption that Janus, who had shared it, publicly corrected; and Alpaca's training data came from text-davinci-003, so Alpaca is a distillation of an RLHF model (B05-28). Sources: Ye et al. · Janus · The Decoder