"The Art of Scaling RL Compute" fits predictable scaling curves to RL training

The paper reports the first large systematic study of RL scaling (over 400,000 GPU-hours), finds that algorithmic recipes have different performance asymptotes while design choices such as loss…

Date
15 October 2025
Who
multi-institution team (authors incl. Rishabh Agarwal, Inderjit Dhillon, David Brandfonbrener)
Confidence
High
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): multi-institution team (authors incl. Rishabh Agarwal, Inderjit Dhillon, David Brandfonbrener) · Confidence: High The paper reports the first large systematic study of RL scaling (over 400,000 GPU-hours), finds that algorithmic recipes have different performance asymptotes while design choices such as loss aggregation, normalisation and curriculum mostly change compute efficiency, and proposes ScaleRL, whose sigmoidal compute-performance curves extrapolated accurately enough to predict a 100,000 GPU-hour run (arXiv 2510.13786). It fills a gap, because by late 2025 the industry bet had shifted to RL compute (o3 at a reported 10x o1's training compute, which Epoch reads as RL compute, per B08-24, and Grok 4's "pretraining-scale" RL) and there was still no public method for predicting RL scaling. Epoch AI's earlier estimate placed R1's RL at roughly 20% of V3's pre-training cost (about $1M against about $5M) and said reasoning-training compute could converge with the frontier within about a year at then-current growth (Epoch, 2025-05); DeepSeek's V3.2 report later said its RL stage exceeded 10% of pre-training cost. See also B03 and B08-25.

Read it in the deep dive