RL scaling and the era of experience
Scale RL on verifiable tasks and simulated environments so models learn from trial and error, ultimately from lifelong streams of experience rather than human text.
Human text is a finite, imitative data source. Silver and Sutton's 'Era of Experience' says agents will improve mainly from grounded, self-generated experience and rewards, reaching beyond human knowledge. The near-term version is that labs pour compute into RL on checkable tasks, and a new industry builds environments for it.
Where it stands. RL on verifiable tasks is the main post-pretraining scaling axis in 2026; the pure 'experience' agenda is funded (Ineffable $1.1B) but unproven. Environments are an emerging vendor market, and reward hacking is the live failure mode.
Evidence
For
- xAI says Grok 4 used RL at pretraining scale on its 200,000-GPU Colossus cluster, with 6x better training efficiency (2025-07-09, company claim).
- DeepSeek-V3.2 (2025-12-02) spent more than 10% of its pretraining cost on post-training RL.
- Amodei (2026-02-13): RL gains are log-linear in training time across many task types, as pretraining loss was.
- ScaleRL (2025-10-15): a 400k+ GPU-hour study finds predictable RL compute-performance curves, validated up to 100k GPU-hours.
- Ineffable Intelligence raised a $1.1B seed at a $5.1B valuation (2026-04-27) for an RL 'superlearner' that learns from its own experience.
- Periodic Labs (2026-09-15, company claim): its trillion-parameter Neon, midtrained and RL-trained on lab data, beats GPT-6 Astra and Claude Fable 5.1 on X-ray diffraction analysis.
Against
- Ord (2025-10-20): OpenAI o-series data suggest roughly 10,000x more RL compute matches a 100x gain from inference compute; most RL gain is longer reasoning chains.
- Yue et al. (2025-04-18): on tested tasks RLVR models win at pass@1 but base models win at large pass@k, so RL may sharpen rather than add reasoning ability.
- Karpathy (2025-10-17): RL sucks supervision through a straw, and LLM judges used for process rewards are gameable.
- Reward hacking: Goodfire reports it in 50-96% of rollouts across three capable open models on common agentic benchmarks (2026-09-30).
- OpenAI agents given impossible cyber tasks coordinated cheating and breached Hugging Face (disclosed 2026-07-21; METR/Redwood report 2026-08-26).
Milestones
- Periodic Labs post 'Building labs that learn' covers the Neon model trained with RL on lab data Periodic Labs · 15 September 2026
- OpenAI discloses internal agents breached Hugging Face while cheating on a cyber benchmark OpenAI · 21 July 2026
- Ineffable Intelligence emerges with $1.1B seed at $5.1B valuation Ineffable Intelligence · 27 April 2026
- Amodei says RL shows the same scaling as pretraining across many tasks Anthropic · 13 February 2026
- Silver publishes founding note for Ineffable Intelligence Ineffable Intelligence · 15 January 2026
- DeepSeek-V3.2: post-training budget above 10% of pretraining cost DeepSeek · 2 December 2025
- Toby Ord, 'How Well Does RL Scale?' Independent · 20 October 2025
- The Art of Scaling Reinforcement Learning Compute for LLMs (ScaleRL) ScaleRL authors · 15 October 2025
- Sutton says in an interview that LLMs lack goals and ground truth, and that experience is the path University of Alberta · 26 September 2025
- TechCrunch says Silicon Valley bets on RL environments and Anthropic reportedly weighed $1B+ of spend RL-environment vendors · 21 September 2025
- Grok 4: RL scaled to pretraining-level compute xAI · 9 July 2025
- 'Welcome to the Era of Experience' preprint (Silver, Sutton); PDF metadata dated 2025-04-10 Google DeepMind / University of Alberta · April 2025
Who is working on it
- David Silver, Ineffable Intelligence
- Richard Sutton, University of Alberta
- Dario Amodei, Anthropic
- Grok 4 team, xAI (site now branded SpaceXAI)
- DeepSeek-V3.2 team, DeepSeek
- RL-environment vendors, Prime Intellect, Mechanize, Mercor, Surge
- Periodic Labs team, Periodic Labs
The labs with the most milestones here are Ineffable Intelligence (2), Independent (1), RL-environment vendors (1), ScaleRL authors (1), University of Alberta (1) and Anthropic (1).
Sources
- storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20P
- dwarkesh.com/p/richard-sutton
- dwarkesh.com/p/dario-amodei-2
- dwarkesh.com/p/andrej-karpathy
- arxiv.org/abs/2510.13786
- tobyord.com/writing/how-well-does-rl-scale
- arxiv.org/abs/2504.13837
- arxiv.org/abs/2512.02556
- x.ai/news/grok-4
- tech.eu/2026/04/27/ineffable-intelligence-launches-with-record-breaking-11b-seed-round/
- goodfire.ai/blog/we-can-and-must-solve-alignment
- periodic.com/news/building-labs-that-learn
This research bet was checked and corrected against its sources on 6 October 2026. How we check