Is scaling over? The 'age of research' debate

Whether more pretraining and RL compute alone gets to AGI, or whether the bottleneck is now new ideas such as generalization and learning from experience.

The thesis tested here is that scaling is no longer sufficient. Sutskever (2025-11-25) says the 2020-25 scaling era is over and generalization is the core gap; Sutton calls LLMs a dead end; LeCun left Meta to build alternatives. Opponents (Amodei, OpenAI) say pretraining and RL curves still hold.

Where it stands. In practice every lab does both, scaling pretraining and RL while betting on one or two research gaps. GPT-6 Astra (Sep 2026) strengthened the short-term 'scaling still works' case, and continual learning and ARC-style generalization stay open.

Evidence

For

  • Sutskever (2025-11-25): 2012-2020 was research, 2020-2025 scaling, 2026 onward research again; models generalize far worse than people despite strong evals.
  • Sutton (2025-09-26): LLMs lack goals and ground truth. Karpathy (2025-10-17): agents are a decade away and RL is a poor supervision channel.
  • LeCun left Meta (reported 2025-11-20) and AMI raised $1.03B (2026-03-09) to build JEPA world models instead.
  • Hassabis (YC interview, May 2026, secondary report): about 50/50 that one or two breakthroughs, such as continual learning, memory and long-term reasoning, are still missing.
  • ARC-AGI-3 launched (2026-03-25) with frontier AI at 0.51% versus humans at 100%; GPT-5.5 scored 0.43% (2026-05-01).
  • Ord (2025-10-20): RL needs roughly 10,000x more compute for the gain 100x inference compute gives.

Against

  • Amodei (2026-02-13): pretraining scaling keeps giving gains and RL shows the same log-linear scaling; a 'country of geniuses' is a few years away.
  • OpenAI's GPT-6 Astra (2026-09-03) came from its largest pretraining run (100,000+ GPUs at Stargate) plus RL; ARC-AGI-3 jumped from 0.43% (GPT-5.5) to 62.7% (ARC-tested semi-private set).
  • Meta's Muse Spark (2026-04-08): a rebuilt stack matches Llama 4 Maverick with over 10x less compute, shown by a fitted pretraining scaling law plus tracked RL and test-time scaling trends.
  • METR time horizons doubled every ~89-131 days (TH 1.1, 2026-01-29).
  • Open-weight labs kept scaling size: Moonshot's Kimi K3 (2026-07-16) has 2.8T parameters and Alibaba's Qwen3.8 (Aug 2026) is 2.4T total, 95B active.
  • Even SSI partnered with Nvidia to raise its compute about 10x to scale its research (2026-07-27): research and scaling are not exclusive.

Milestones

Who is working on it

  • Ilya Sutskever (age of research), SSI
  • Richard Sutton (LLMs a dead end), University of Alberta
  • Yann LeCun (needs world models), AMI Labs
  • Andrej Karpathy (decade of agents), Eureka Labs
  • Gary Marcus (scaling is flattening), Marcus on AI
  • Demis Hassabis (about 50/50: scale plus one or two breakthroughs), Google DeepMind
  • Dario Amodei (scaling still holds), Anthropic
  • Greg Brockman, Aidan Clark (largest training run), OpenAI

The labs with the most milestones here are Safe Superintelligence Inc. (SSI) (2), ARC Prize Foundation (1), Marcus on AI (1), University of Alberta (1), Anthropic (1) and Eureka Labs (1).

Sources

  1. dwarkesh.com/p/ilya-sutskever-2
  2. dwarkesh.com/p/richard-sutton
  3. dwarkesh.com/p/andrej-karpathy
  4. dwarkesh.com/p/dario-amodei-2
  5. garymarcus.substack.com/p/a-trillion-dollars-is-a-terrible
  6. the-decoder.com/gpt-6-astra-is-the-first-model-making-openai-willing-to-declare-the-agi-er
  7. fortune.com/2026/09/03/openai-debuts-gpt-6-astra-computer-use-greg-brockman-says-start-of-
  8. arcprize.org/blog/astra
  9. ai.meta.com/blog/introducing-muse-spark-msl/
  10. metr.org/blog/2026-1-29-time-horizon-1-1/
  11. kimi.com/blog/kimi-k3
  12. huggingface.co/Qwen/Qwen3.8-2.4T-A95B
  13. eweek.com/news/nvidia-ssi-5b-ai-partnership/
  14. the-ai-corner.com/p/demis-hassabis-agi-2030-deep-tech-founder-playbook-2026
  15. tobyord.com/writing/how-well-does-rl-scale
  16. techcrunch.com/2026/03/09/yann-lecuns-ami-labs-raises-1-03-billion-to-build-world-models/
  17. the-decoder.com/yann-lecun-leaves-meta-to-launch-new-ai-startup/
  18. arcprize.org/blog/arc-agi-3-launch
  19. arcprize.org/blog/arc-agi-3-gpt-5-5-opus-4-7-analysis
  20. evertiq.com/news/2026-07-28-nvidia-to-invest-5b-in-startup-accesses-its-closely-guarded-re
  21. fortune.com/2026/07/16/moonshots-kimi-k3-pushes-chinese-ai-into-fable-level-territory/

This research bet was checked and corrected against its sources on 6 October 2026. How we check