Moonshot published its Kimi k1.5 reinforcement learning recipe on the same day as R1
Moonshot announced Kimi k1.5 the same day as R1 (reported 2025-01-21 by The Decoder; arXiv v1 2025-01-22).
- Date
- 20 January 2025
- Who
- Moonshot AI
- Confidence
- High
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): Moonshot AI · Confidence: High Moonshot announced Kimi k1.5 the same day as R1 (reported 2025-01-21 by The Decoder; arXiv v1 2025-01-22). The report described a deliberately simple RL framework that scales the RL context window to 128K tokens (using "partial rollouts" that carry long responses across training iterations), uses a variant of online mirror descent for policy optimisation, and explicitly avoids Monte Carlo tree search, value functions and process reward models (arXiv 2501.12599). It reported AIME 77.5, MATH-500 96.2, Codeforces 94th percentile and MathVista 74.9, plus a "long2short" method that distils long-CoT into short-CoT (AIME 60.8 for the short model). I found no weight release for k1.5, and the project repository carries the report and benchmark charts only (GitHub), which is why this entry is marked "no weights found." I infer that two Chinese labs, apparently independently, reported the same "simple outcome-reward RL plus long context" conclusion within hours of each other, which suggests that any strong lab could find the recipe. Moonshot's later line is in B08-41.