s1 fine-tuned Qwen2.5-32B on 1,000 examples and extended its thinking by appending "Wait"

s1 fine-tuned Qwen2.5-32B-Instruct on only 1,000 reasoning examples ("s1K," selected from 59,029 candidates for difficulty, diversity and quality, with traces generated via Google's Gemini Flash…

Date
31 January 2025
Who
Stanford, University of Washington, Allen Institute for AI, Contextual AI
People
Niklas Muennighoff, Percy Liang, Emmanuel Candès, Tatsunori Hashimoto, Li Fei-Fei
Confidence
High
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 4/5 · Org(s): Stanford, University of Washington, Allen Institute for AI, Contextual AI · People: Niklas Muennighoff, Percy Liang, Emmanuel Candès, Tatsunori Hashimoto, Li Fei-Fei · Confidence: High s1 fine-tuned Qwen2.5-32B-Instruct on only 1,000 reasoning examples ("s1K," selected from 59,029 candidates for difficulty, diversity and quality, with traces generated via Google's Gemini Flash Thinking API) in 26 minutes on 16 H100 GPUs, then controlled inference compute with "budget forcing", which cuts thinking off or appends "Wait" when the model tries to stop so it re-checks its answer. Accuracy on AIME 2024 rose from 50% to 57% as thinking was extended, and the authors report exceeding o1-preview on competition maths by up to 27% (arXiv 2501.19393, v1 2025-01-31; PDF read locally). It mattered as a minimal, fully open demonstration that long-CoT behaviour can be elicited cheaply from a strong base and that test-time compute can be controlled. The caveat is that the teacher traces came from a proprietary reasoning model, so it is a case of distillation. Connect to B07 (test-time compute theory) and B08-25.

Read it in the deep dive