Google's Gemini 1.0 report documents SFT, reward modeling and RLHF

Google's Gemini report describes post-training in four stages. The first collects diverse prompts, the second applies supervised fine-tuning on demonstrations (human-written or model-generated and…

Date
6 December 2023
Who
Google DeepMind
Confidence
High (Google's report)
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): Google DeepMind · Confidence: High (Google's report) Google's Gemini report describes post-training in four stages. The first collects diverse prompts, the second applies supervised fine-tuning on demonstrations (human-written or model-generated and reviewed), the third trains reward models on human feedback data, such as relative preferences, and the fourth applies RLHF. Data sources include vendor-created data, third-party licensed sources and synthetic approaches. It reports that RLHF improved multimodal tasks, with a side-by-side score gain of +0.223 (±0.06) on image understanding for a Gemini Apps Pro model with SFT and RLHF versus SFT alone (Gemini report §6.3; announced 2023-12-06 per Google; arXiv v1 2023-12-19). Lag from InstructGPT's blog (2022-01-27) to this published description is about 22 months, though Google's own work (Sparrow, B05-16) used RLHF earlier. Sources: Gemini Team · Google

Read it in the deep dive