Berkeley and Hugging Face started open reproductions of reasoning models with Sky-T1, TinyZero and Open-R1
Two routes to "your own reasoning model" appeared within weeks. Distillation-SFT: Berkeley's Sky-T1-32B-Preview (2025-01-10, before R1) fine-tuned Qwen2.5-32B-Instruct on 17K traces distilled from…
- Date
- 28 January 2025
- Who
- UC Berkeley (NovaSky / Sky Computing Lab; Jiayi Pan's team), Hugging Face
- People
- Jiayi Pan, Leandro von Werra, Lewis Tunstall, Elie Bakouch
- Confidence
- High
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 4/5 · Org(s): UC Berkeley (NovaSky / Sky Computing Lab; Jiayi Pan's team), Hugging Face · People: Jiayi Pan, Leandro von Werra, Lewis Tunstall, Elie Bakouch · Confidence: High Two routes to "your own reasoning model" appeared within weeks. Distillation-SFT: Berkeley's Sky-T1-32B-Preview (2025-01-10, before R1) fine-tuned Qwen2.5-32B-Instruct on 17K traces distilled from QwQ-32B-Preview, claiming o1-preview-level scores on some benchmarks for under $450 of compute (AIME 2024 43.3%, MATH-500 82.4%) (NovaSky post); s1 did the same with 1,000 examples. Small-scale RL: TinyZero (about 2025-01-24) reproduced R1-Zero's "aha" on a Countdown arithmetic task with a 3B Qwen model for under $30 (repo), and Hugging Face's Open-R1 (2025-01-28) laid out a plan to rebuild what DeepSeek withheld, namely the R1-distilled datasets, the pure-RL pipeline and the multi-stage recipe (blog). I infer that together they showed reasoning behaviour appears at toy scale while frontier-level results need the large base model and an industrial RL system, a gap the DAPO paper later tried to close.