xAI releases Grok 4 and Grok 4 Heavy, with RL at pretraining scale and parallel test-time compute
xAI said Grok 4's reasoning was refined with reinforcement learning "at pretraining scale," using more than an order of magnitude more RL compute than earlier models, natively trained with tools.
- Date
- 9 July 2025
- Who
- xAI
- Confidence
- Medium (company-reported)
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): xAI · Confidence: Medium (company-reported) xAI said Grok 4's reasoning was refined with reinforcement learning "at pretraining scale," using more than an order of magnitude more RL compute than earlier models, natively trained with tools. Grok 4 Heavy used parallel test-time compute with multiple agents and was reported by xAI as the first to score 50.7% on the text-only subset of Humanity's Last Exam (B08-11a). xAI attributes the 15.9% ARC-AGI-2 result, a closed-model record at the time, to the single-agent Grok 4 and not to Heavy (x.ai). It is an early public statement of the "scale RL like pretraining" thesis, ahead of DeepSeek's V3.2 paper (B08-43); the Humanity's Last Exam figures are company-measured and were with tools for some variants. It came 170 days after R1.