Is scaling over? The 'age of research' debate
Whether more pretraining and RL compute alone gets to AGI, or whether the bottleneck is now new ideas such as generalization and learning from experience.
The thesis tested here is that scaling is no longer sufficient. Sutskever (2025-11-25) says the 2020-25 scaling era is over and generalization is the core gap; Sutton calls LLMs a dead end; LeCun left Meta to build alternatives. Opponents (Amodei, OpenAI) say pretraining and RL curves still hold.
Where it stands. In practice every lab does both, scaling pretraining and RL while betting on one or two research gaps. GPT-6 Astra (Sep 2026) strengthened the short-term 'scaling still works' case, and continual learning and ARC-style generalization stay open.
Evidence
For
- Sutskever (2025-11-25): 2012-2020 was research, 2020-2025 scaling, 2026 onward research again; models generalize far worse than people despite strong evals.
- Sutton (2025-09-26): LLMs lack goals and ground truth. Karpathy (2025-10-17): agents are a decade away and RL is a poor supervision channel.
- LeCun left Meta (reported 2025-11-20) and AMI raised $1.03B (2026-03-09) to build JEPA world models instead.
- Hassabis (YC interview, May 2026, secondary report): about 50/50 that one or two breakthroughs, such as continual learning, memory and long-term reasoning, are still missing.
- ARC-AGI-3 launched (2026-03-25) with frontier AI at 0.51% versus humans at 100%; GPT-5.5 scored 0.43% (2026-05-01).
- Ord (2025-10-20): RL needs roughly 10,000x more compute for the gain 100x inference compute gives.
Against
- Amodei (2026-02-13): pretraining scaling keeps giving gains and RL shows the same log-linear scaling; a 'country of geniuses' is a few years away.
- OpenAI's GPT-6 Astra (2026-09-03) came from its largest pretraining run (100,000+ GPUs at Stargate) plus RL; ARC-AGI-3 jumped from 0.43% (GPT-5.5) to 62.7% (ARC-tested semi-private set).
- Meta's Muse Spark (2026-04-08): a rebuilt stack matches Llama 4 Maverick with over 10x less compute, shown by a fitted pretraining scaling law plus tracked RL and test-time scaling trends.
- METR time horizons doubled every ~89-131 days (TH 1.1, 2026-01-29).
- Open-weight labs kept scaling size: Moonshot's Kimi K3 (2026-07-16) has 2.8T parameters and Alibaba's Qwen3.8 (Aug 2026) is 2.4T total, 95B active.
- Even SSI partnered with Nvidia to raise its compute about 10x to scale its research (2026-07-27): research and scaling are not exclusive.
Milestones
- GPT-6 Astra is OpenAI's largest pretraining run and scores 62.7% on ARC-AGI-3 with the standard harness OpenAI · 3 September 2026
- Nvidia-SSI partnership provides 10x compute for research 'worthy of scaling up' SSI / NVIDIA · 27 July 2026
- Kimi K3: 2.8T-parameter open model, first open 3T-class (company claim) Moonshot AI · 16 July 2026
- Hassabis YC interview, as summarized in a 2026-05-12 article, puts AGI at about 2030 and is 50/50 on missing breakthroughs Google DeepMind · 12 May 2026
- Meta Muse Spark reaches >10x compute efficiency vs Llama 4 Maverick on a rebuilt stack Meta · 8 April 2026
- ARC-AGI-3 launched; frontier AI 0.51% vs humans 100% ARC Prize Foundation · 25 March 2026
- Amodei says 'we are near the end of the exponential', that all scaling holds and that RL is log-linear Anthropic · 13 February 2026
- Marcus reads Sutskever's interview as scaling flattening out Marcus on AI · 27 November 2025
- Sutskever says we are moving from the age of scaling to the age of research SSI · 25 November 2025
- Karpathy covers 'decade of agents', RL criticisms and missing cognitive components Eureka Labs · 17 October 2025
- Sutton says LLMs are a dead end and backs experiential learning instead University of Alberta · 26 September 2025
Who is working on it
- Ilya Sutskever (age of research), SSI
- Richard Sutton (LLMs a dead end), University of Alberta
- Yann LeCun (needs world models), AMI Labs
- Andrej Karpathy (decade of agents), Eureka Labs
- Gary Marcus (scaling is flattening), Marcus on AI
- Demis Hassabis (about 50/50: scale plus one or two breakthroughs), Google DeepMind
- Dario Amodei (scaling still holds), Anthropic
- Greg Brockman, Aidan Clark (largest training run), OpenAI
The labs with the most milestones here are Safe Superintelligence Inc. (SSI) (2), ARC Prize Foundation (1), Marcus on AI (1), University of Alberta (1), Anthropic (1) and Eureka Labs (1).
Sources
- dwarkesh.com/p/ilya-sutskever-2
- dwarkesh.com/p/richard-sutton
- dwarkesh.com/p/andrej-karpathy
- dwarkesh.com/p/dario-amodei-2
- garymarcus.substack.com/p/a-trillion-dollars-is-a-terrible
- the-decoder.com/gpt-6-astra-is-the-first-model-making-openai-willing-to-declare-the-agi-er
- fortune.com/2026/09/03/openai-debuts-gpt-6-astra-computer-use-greg-brockman-says-start-of-
- arcprize.org/blog/astra
- ai.meta.com/blog/introducing-muse-spark-msl/
- metr.org/blog/2026-1-29-time-horizon-1-1/
- kimi.com/blog/kimi-k3
- huggingface.co/Qwen/Qwen3.8-2.4T-A95B
- eweek.com/news/nvidia-ssi-5b-ai-partnership/
- the-ai-corner.com/p/demis-hassabis-agi-2030-deep-tech-founder-playbook-2026
- tobyord.com/writing/how-well-does-rl-scale
- techcrunch.com/2026/03/09/yann-lecuns-ami-labs-raises-1-03-billion-to-build-world-models/
- the-decoder.com/yann-lecun-leaves-meta-to-launch-new-ai-startup/
- arcprize.org/blog/arc-agi-3-launch
- arcprize.org/blog/arc-agi-3-gpt-5-5-opus-4-7-analysis
- evertiq.com/news/2026-07-28-nvidia-to-invest-5b-in-startup-accesses-its-closely-guarded-re
- fortune.com/2026/07/16/moonshots-kimi-k3-pushes-chinese-ai-into-fable-level-territory/
This research bet was checked and corrected against its sources on 6 October 2026. How we check