Continual learning and memory
Systems that keep learning after deployment, instead of staying frozen at a training cutoff, through weight updates, test-time training or external memory.
Frozen models generalize worse than people and cannot learn on the job; this, not raw scale, is the next bottleneck. Sutskever imagines a deployed 'superintelligent 15-year-old' that learns from experience; Karpathy wants a sleep-like distillation phase; Google proposes nested, multi-speed learning.
Where it stands. Widely named the top missing capability, with no agreed method. Anthropic expects a fix within one to two years, and DeepMind says it is uncracked. No lab was found shipping online weight updates, and memory and skills layers fill the gap.
Evidence
For
- Dwarkesh Patel (2025-06-02) called missing continual learning a huge bottleneck; Karpathy (2025-10-17) and Sutskever (2025-11-25) independently cite frozen weights and poor generalization.
- Architectures: Titans (2024-12-31) adds a neural memory updated at test time; Nested Learning/Hope (Google, 2025-11-07); TTT-E2E (2025-12-29) compresses context into weights at test time.
- TTT-E2E: a 3B model trained on 164B tokens scales with context length like full attention, at constant latency and 2.7x faster at 128K (authors' result).
- Anthropic's Sholto Douglas predicted continual learning gets solved in 2026; Amodei (2026-02-13) sees a good chance of solving it within one to two years.
- SSI's Nvidia partnership (2026-07-27): Sutskever says SSI has research worth scaling; Nvidia reportedly invests $5B and adds 10x compute.
Against
- Amodei argues continual learning may not be a barrier at all: pretraining and RL generalization, or very long in-context learning, could deliver on-the-job skill (2026-02-13).
- Not found: any lab shipping online weight updates in a flagship model by 2026-10-04; practice is memory files, skills, retrieval and LoRA adapters.
- Letta (2025-12-11) argues updating learned context in token space, not weights, is the practical primitive for agents.
- Catastrophic forgetting persists at scale, and updating weights from user feedback invites manipulation and could invalidate earlier safety tests (Transformer News, 2026-01-22).
- Hassabis (YC interview, May 2026, secondary report) calls million-token context 'duct tape' and says the problem is not yet cracked.
- Product answer so far is memory, not learning: OpenAI shipped a new ChatGPT memory system, 'Dreaming' (2026-06-04), to keep remembered preferences fresh across chats.
Milestones
- Nvidia-SSI partnership with a reported $5B investment, and SSI says its research is worth scaling SSI / NVIDIA · 27 July 2026
- Thinking Machines says most AI is trained once and frozen, and it is pursuing customizable models and weight training Thinking Machines Lab · 10 July 2026
- OpenAI ships 'Dreaming', a new ChatGPT memory system for remembering preferences across conversations OpenAI · 4 June 2026
- Hassabis YC interview, as summarized in a 2026-05-12 article, says continual learning and memory are still missing Google DeepMind · 12 May 2026
- Amodei says there is a good chance continual learning is solved in 1-2 years, but it may not be necessary Anthropic · 13 February 2026
- End-to-End Test-Time Training for Long Context (arXiv v1) Stanford, NVIDIA and collaborators · 29 December 2025
- Letta puts continual learning in token space and leaves the weights alone Letta · 11 December 2025
- Sutskever says the field is moving from the age of scaling to the age of research and describes deployed continual learners SSI · 25 November 2025
- Nested Learning and the Hope architecture introduced Google Research · 7 November 2025
- Karpathy calls it the 'decade of agents' and says models need a distillation phase to consolidate learning Eureka Labs · 17 October 2025
- Dwarkesh Patel essay names continual learning a huge bottleneck to AGI Dwarkesh Podcast · 2 June 2025
- Titans: Learning to Memorize at Test Time (arXiv v1) Google Research · 31 December 2024
Who is working on it
- Ilya Sutskever, Safe Superintelligence Inc. (SSI)
- Andrej Karpathy, Eureka Labs
- Dwarkesh Patel (framing the bottleneck), Dwarkesh Podcast
- Ali Behrouz, Vahab Mirrokni, Google Research
- Yu Sun and co-authors (TTT-E2E), Stanford, NVIDIA and collaborators
- Demis Hassabis, Google DeepMind
- Dario Amodei, Sholto Douglas, Anthropic
- Thinking Machines team, Thinking Machines Lab
- Letta team (token-space learning), Letta
The labs with the most milestones here are Google DeepMind (3), Safe Superintelligence Inc. (SSI) (2), Dwarkesh Podcast (1), Letta (1), Stanford, NVIDIA and collaborators (1) and Anthropic (1).
Sources
- dwarkesh.com/p/ilya-sutskever-2
- dwarkesh.com/p/andrej-karpathy
- dwarkesh.com/p/dario-amodei-2
- dwarkesh.com/p/timelines-june-2025
- research.google/blog/introducing-nested-learning-a-new-ml-paradigm-for-continual-learning/
- arxiv.org/abs/2512.23675
- transformernews.ai/p/teaching-ai-to-continual-learning
- letta.com/blog/continual-learning/
- the-ai-corner.com/p/demis-hassabis-agi-2030-deep-tech-founder-playbook-2026
- openai.com/news/rss.xml
- thestar.com.my/tech/tech-news/2026/07/28/nvidia-to-invest-5-billion-in-ilya-sutskever039s-
This research bet was checked and corrected against its sources on 6 October 2026. How we check