Google publishes LaMDA, a dialogue model fine-tuned without RL
LaMDA was a family of Transformer dialogue models up to 137B parameters, pretrained on 1.56T words of public dialogue and web text.
- Date
- 20 January 2022
- Who
- People
- Romal Thoppilan, Daniel De Freitas, Noam Shazeer and others (60 authors)
- Confidence
- High (paper); Medium (reporting on release decisions)
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): Google · People: Romal Thoppilan, Daniel De Freitas, Noam Shazeer and others (60 authors) · Confidence: High (paper); Medium (reporting on release decisions) LaMDA was a family of Transformer dialogue models up to 137B parameters, pretrained on 1.56T words of public dialogue and web text. Google improved it with crowdworker-annotated data for quality, safety and groundedness, used to fine-tune discriminators that filter and re-rank candidate responses, plus a tool for consulting external sources; there is no reinforcement learning stage (arXiv:2201.08239, v1 2022-01-20; DeepMind's Sparrow paper makes the same observation and notes that LaMDA uses supervised learning and ranking, with no RL, Sparrow).
Google had unveiled LaMDA two years before its 2023-02-06 Bard post (Google) but did not release a chatbot. The Wall Street Journal reported (as summarized by Yahoo) that executives blocked attempts by its builders, Daniel De Freitas and Noam Shazeer, to share the model with outside researchers, add it to Google Assistant or demo it publicly on safety and fairness grounds, and that both left near the end of 2021 to start Character.AI (Yahoo summary of WSJ). On 2022-07-22 Google fired Blake Lemoine, the engineer who publicly claimed LaMDA was sentient (Big Technology). After ChatGPT, Sundar Pichai and Jeff Dean told an all-hands that Google had similar capabilities but more reputational risk if answers were wrong, and Dean said Google was moving "more conservatively than a small startup" (CNBC, 2022-12-13).
Inference: Alphabet had both halves, but in separate organizations. LaMDA (and later PaLM and Flan) sat at Google Brain, which paired a base model with a classifier-and-filtering stack and no RL stage, and RLHF dialogue agents sat at DeepMind (GopherCite, 2022-03-21, and Sparrow, 2022-09-22; B05-13a, B05-16), and neither shipped as a product. Product and brand constraints differed, and the two groups merged only on 2023-04-20 (B05-27a). Connects to B05-24. Sources: Thoppilan et al. · CNBC