Continual learning and memory

Systems that keep learning after deployment, instead of staying frozen at a training cutoff, through weight updates, test-time training or external memory.

Frozen models generalize worse than people and cannot learn on the job; this, not raw scale, is the next bottleneck. Sutskever imagines a deployed 'superintelligent 15-year-old' that learns from experience; Karpathy wants a sleep-like distillation phase; Google proposes nested, multi-speed learning.

Where it stands. Widely named the top missing capability, with no agreed method. Anthropic expects a fix within one to two years, and DeepMind says it is uncracked. No lab was found shipping online weight updates, and memory and skills layers fill the gap.

Evidence

For

  • Dwarkesh Patel (2025-06-02) called missing continual learning a huge bottleneck; Karpathy (2025-10-17) and Sutskever (2025-11-25) independently cite frozen weights and poor generalization.
  • Architectures: Titans (2024-12-31) adds a neural memory updated at test time; Nested Learning/Hope (Google, 2025-11-07); TTT-E2E (2025-12-29) compresses context into weights at test time.
  • TTT-E2E: a 3B model trained on 164B tokens scales with context length like full attention, at constant latency and 2.7x faster at 128K (authors' result).
  • Anthropic's Sholto Douglas predicted continual learning gets solved in 2026; Amodei (2026-02-13) sees a good chance of solving it within one to two years.
  • SSI's Nvidia partnership (2026-07-27): Sutskever says SSI has research worth scaling; Nvidia reportedly invests $5B and adds 10x compute.

Against

  • Amodei argues continual learning may not be a barrier at all: pretraining and RL generalization, or very long in-context learning, could deliver on-the-job skill (2026-02-13).
  • Not found: any lab shipping online weight updates in a flagship model by 2026-10-04; practice is memory files, skills, retrieval and LoRA adapters.
  • Letta (2025-12-11) argues updating learned context in token space, not weights, is the practical primitive for agents.
  • Catastrophic forgetting persists at scale, and updating weights from user feedback invites manipulation and could invalidate earlier safety tests (Transformer News, 2026-01-22).
  • Hassabis (YC interview, May 2026, secondary report) calls million-token context 'duct tape' and says the problem is not yet cracked.
  • Product answer so far is memory, not learning: OpenAI shipped a new ChatGPT memory system, 'Dreaming' (2026-06-04), to keep remembered preferences fresh across chats.

Milestones

Who is working on it

The labs with the most milestones here are Google DeepMind (3), Safe Superintelligence Inc. (SSI) (2), Dwarkesh Podcast (1), Letta (1), Stanford, NVIDIA and collaborators (1) and Anthropic (1).

Sources

  1. dwarkesh.com/p/ilya-sutskever-2
  2. dwarkesh.com/p/andrej-karpathy
  3. dwarkesh.com/p/dario-amodei-2
  4. dwarkesh.com/p/timelines-june-2025
  5. research.google/blog/introducing-nested-learning-a-new-ml-paradigm-for-continual-learning/
  6. arxiv.org/abs/2512.23675
  7. transformernews.ai/p/teaching-ai-to-continual-learning
  8. letta.com/blog/continual-learning/
  9. the-ai-corner.com/p/demis-hassabis-agi-2030-deep-tech-founder-playbook-2026
  10. openai.com/news/rss.xml
  11. thestar.com.my/tech/tech-news/2026/07/28/nvidia-to-invest-5-billion-in-ilya-sutskever039s-

This research bet was checked and corrected against its sources on 6 October 2026. How we check