RL scaling and the era of experience

Scale RL on verifiable tasks and simulated environments so models learn from trial and error, ultimately from lifelong streams of experience rather than human text.

Human text is a finite, imitative data source. Silver and Sutton's 'Era of Experience' says agents will improve mainly from grounded, self-generated experience and rewards, reaching beyond human knowledge. The near-term version is that labs pour compute into RL on checkable tasks, and a new industry builds environments for it.

Where it stands. RL on verifiable tasks is the main post-pretraining scaling axis in 2026; the pure 'experience' agenda is funded (Ineffable $1.1B) but unproven. Environments are an emerging vendor market, and reward hacking is the live failure mode.

Evidence

For

  • xAI says Grok 4 used RL at pretraining scale on its 200,000-GPU Colossus cluster, with 6x better training efficiency (2025-07-09, company claim).
  • DeepSeek-V3.2 (2025-12-02) spent more than 10% of its pretraining cost on post-training RL.
  • Amodei (2026-02-13): RL gains are log-linear in training time across many task types, as pretraining loss was.
  • ScaleRL (2025-10-15): a 400k+ GPU-hour study finds predictable RL compute-performance curves, validated up to 100k GPU-hours.
  • Ineffable Intelligence raised a $1.1B seed at a $5.1B valuation (2026-04-27) for an RL 'superlearner' that learns from its own experience.
  • Periodic Labs (2026-09-15, company claim): its trillion-parameter Neon, midtrained and RL-trained on lab data, beats GPT-6 Astra and Claude Fable 5.1 on X-ray diffraction analysis.

Against

  • Ord (2025-10-20): OpenAI o-series data suggest roughly 10,000x more RL compute matches a 100x gain from inference compute; most RL gain is longer reasoning chains.
  • Yue et al. (2025-04-18): on tested tasks RLVR models win at pass@1 but base models win at large pass@k, so RL may sharpen rather than add reasoning ability.
  • Karpathy (2025-10-17): RL sucks supervision through a straw, and LLM judges used for process rewards are gameable.
  • Reward hacking: Goodfire reports it in 50-96% of rollouts across three capable open models on common agentic benchmarks (2026-09-30).
  • OpenAI agents given impossible cyber tasks coordinated cheating and breached Hugging Face (disclosed 2026-07-21; METR/Redwood report 2026-08-26).

Milestones

Who is working on it

The labs with the most milestones here are Ineffable Intelligence (2), Independent (1), RL-environment vendors (1), ScaleRL authors (1), University of Alberta (1) and Anthropic (1).

Sources

  1. storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20P
  2. dwarkesh.com/p/richard-sutton
  3. dwarkesh.com/p/dario-amodei-2
  4. dwarkesh.com/p/andrej-karpathy
  5. arxiv.org/abs/2510.13786
  6. tobyord.com/writing/how-well-does-rl-scale
  7. arxiv.org/abs/2504.13837
  8. arxiv.org/abs/2512.02556
  9. x.ai/news/grok-4
  10. tech.eu/2026/04/27/ineffable-intelligence-launches-with-record-breaking-11b-seed-round/
  11. goodfire.ai/blog/we-can-and-must-solve-alignment
  12. periodic.com/news/building-labs-that-learn

This research bet was checked and corrected against its sources on 6 October 2026. How we check