MiniMax-M1 pairs hybrid attention with a $535K RL run

MiniMax-M1 (456B total, 45.9B active, 1M-token context, open weights) pairs a lightning-attention hybrid mixture-of-experts with a new RL algorithm, CISPO, which clips importance-sampling weights and…

Date
16 June 2025
Who
MiniMax
Confidence
High
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 2/5 · Org(s): MiniMax · Confidence: High MiniMax-M1 (456B total, 45.9B active, 1M-token context, open weights) pairs a lightning-attention hybrid mixture-of-experts with a new RL algorithm, CISPO, which clips importance-sampling weights and does not clip token updates; the report says full RL took three weeks on 512 H800s at about $534,700 of rental cost, and releases 40K and 80K thinking-budget variants (arXiv 2506.13585). It is an early public dollar figure for a reasoning RL run, three months before DeepSeek's own $294K figure for R1 appeared (B08-15); the authors present efficient attention as what makes long-thinking RL affordable (B02). For comparison, the later V3.2 statement said that RL exceeded 10% of pre-training cost.

Read it in the deep dive