OpenAI trains WebGPT to browse the web and answer questions from human feedback

WebGPT taught GPT-3 (760M/13B/175B) to answer ELI5 questions by browsing the web through the Bing API.

Date
17 December 2021
Who
OpenAI
People
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, John Schulman and others
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI · People: Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, John Schulman and others · Confidence: High WebGPT taught GPT-3 (760M/13B/175B) to answer ELI5 questions by browsing the web through the Bing API. Training started with behavior cloning on about 6,000 human demonstrations of the browsing environment, then added a reward model (about 21,500 comparisons were collected, of which around 16,000 trained the final reward models and the rest were held out for validation), then rejection sampling (best-of-n) and some RL. The best 175B best-of-64 model was preferred to human demonstrators 56% of the time and to the top Reddit answer 69% (arXiv:2112.09332, v1 2021-12-17).

It belongs here because in a Berkeley talk on 2023-04-19 John Schulman used this work as the example of RLHF for factuality and noted that despite elaborate annotation interfaces the labeling signal collapsed to a single bit of preference and that labelers may have been swayed by confident, citation-laden style (B05-32; transcript, published 2023-04-24); OpenAI's ChatGPT post attributes factuality limits to the lack of a source of truth in RL (ChatGPT post). It also made the model collect references while browsing so that human evaluation of factual accuracy was easier; DeepMind's Sparrow uses a similar evidence-in-the-loop design (B05-16). Sources: Nakano et al.

Read it in the deep dive