text-davinci-001/002/003 and the GPT-3.5 series, and which of them used RLHF

OpenAI's API model names did not say which models were trained with RLHF. In the lineage described by OpenAI's model index (the page itself is no longer retrievable; I relied on a Fudan University…

Date
28 November 2022
Who
OpenAI
Confidence
High (for the method labels); Medium (exact release days)
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · Confidence: High (for the method labels); Medium (exact release days) OpenAI's API model names did not say which models were trained with RLHF. In the lineage described by OpenAI's model index (the page itself is no longer retrievable; I relied on a Fudan University analysis that reproduces it and on a contemporaneous LessWrong thread), the order is base davinci (2020); text-davinci-001, an InstructGPT model trained with "FeedME", meaning supervised fine-tuning on human demonstrations plus model samples that labelers rated highly (Jan Leike told the LessWrong poster the sample-generating models were a mix of earlier ones); code-davinci-002, the code model that is the base of the GPT-3.5 series; text-davinci-002, FeedME on that base; text-davinci-003, trained with PPO against a reward model trained from human comparisons; then gpt-3.5-turbo, the chat-optimized model (Ye et al., arXiv:2303.10420; LessWrong/Janus update, 2022-11-19, with 2022-11-30 comments). The InstructGPT blog had already noted that the deployed API models used a similar but slightly different method than the paper's (blog footnote). text-davinci-003 was announced on Monday 2022-11-28; Jan Leike described it as largely equivalent to the InstructGPT models but not identical, scoring higher on human preference without being more capable (The Decoder, 2022-11-29; VentureBeat). Why it matters for interviews: "text-davinci-002 is RLHF" was a widespread community assumption that Janus, who had shared it, publicly corrected; and Alpaca's training data came from text-davinci-003, so Alpaca is a distillation of an RLHF model (B05-28). Sources: Ye et al. · Janus · The Decoder

Read it in the deep dive