OpenAI trains WebGPT to browse the web and answer questions from human feedback
WebGPT taught GPT-3 (760M/13B/175B) to answer ELI5 questions by browsing the web through the Bing API.
- Date
- 17 December 2021
- Who
- OpenAI
- People
- Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, John Schulman and others
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · People: Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, John Schulman and others · Confidence: High WebGPT taught GPT-3 (760M/13B/175B) to answer ELI5 questions by browsing the web through the Bing API. Training started with behavior cloning on about 6,000 human demonstrations of the browsing environment, then added a reward model (about 21,500 comparisons were collected, of which around 16,000 trained the final reward models and the rest were held out for validation), then rejection sampling (best-of-n) and some RL. The best 175B best-of-64 model was preferred to human demonstrators 56% of the time and to the top Reddit answer 69% (arXiv:2112.09332, v1 2021-12-17).
It belongs here because in a Berkeley talk on 2023-04-19 John Schulman used this work as the example of RLHF for factuality and noted that despite elaborate annotation interfaces the labeling signal collapsed to a single bit of preference and that labelers may have been swayed by confident, citation-laden style (B05-32; transcript, published 2023-04-24); OpenAI's ChatGPT post attributes factuality limits to the lack of a source of truth in RL (ChatGPT post). It also made the model collect references while browsing so that human evaluation of factual accuracy was easier; DeepMind's Sparrow uses a similar evidence-in-the-loop design (B05-16). Sources: Nakano et al.