Meta releases BlenderBot 3 and later pauses the Galactica demo

Meta released BlenderBot 3, a 175B conversational agent built from OPT-175B that learned from feedback by people chatting with it, and warned that it could still make rude or offensive comments…

Date
5 August 2022
Who
Meta AI
Confidence
High (dates); Medium (causal reading)
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 2/5 · Org(s): Meta AI · Confidence: High (dates); Medium (causal reading) Meta released BlenderBot 3, a 175B conversational agent built from OPT-175B that learned from feedback by people chatting with it, and warned that it could still make rude or offensive comments (Meta, 2022-08-05; arXiv:2208.03188). On 2022-11-15 it released Galactica, a science-focused language model; after three days of criticism over fabricated papers and confident falsehoods, Meta paused the public demo (MIT Technology Review, 2022-11-18; arXiv:2211.09085). The New York Times later reported that some OpenAI employees doubted a chatbot would succeed because BlenderBot had flopped and Galactica was pulled (NYT via Khaleej Times). Neither paper describes an InstructGPT-style PPO RLHF stage (BlenderBot 3 learns from user feedback signals; Galactica is a pretrained model); ChatGPT launched two weeks after Galactica. Inference from author lists: eight of Galactica's nine authors (including Thomas Scialom, Ross Taylor and Robert Stojnic) also appear on the Llama 2 paper, whose corresponding authors are Scialom and Hugo Touvron, so Meta's post-ChatGPT RLHF capability was assembled partly from the team that had shipped, and retreated from, Galactica (Llama 2 author list). Sources: above.

Read it in the deep dive