DeepSeek-R1-0528 lifts AIME 2025 from 70.0% to 87.5% with more RL compute
A mid-life update to R1 with the same architecture and API. DeepSeek credits "increased computational resources and algorithmic optimization mechanisms during post-training" for the gains.
- Date
- 28 May 2025
- Who
- DeepSeek
- Confidence
- High
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): DeepSeek · Confidence: High A mid-life update to R1 with the same architecture and API. DeepSeek credits "increased computational resources and algorithmic optimization mechanisms during post-training" for the gains. AIME 2025 rose from 70.0% to 87.5%, GPQA Diamond from 71.5% to 81.0%, Humanity's Last Exam from 8.5% to 17.7%, and LiveCodeBench from 63.5% to 73.3%, while average tokens per AIME question roughly doubled from 12K to 23K (model card). The same card releases a distilled Qwen3-8B reaching 86.0% on AIME 2024. It is the cleanest public data point for the claim that RL compute and thinking length both buy accuracy, and it landed nearly three months (85 days) before the V3.1 hybrid (B08-37). Later evaluations such as NIST's used this version as DeepSeek's "most secure" model (B08-12).