Next bets

The research directions labs are betting on next, and the evidence for each.

  • World models
    Learn a predictive model of how the world responds to actions and plan inside it. Rivals disagree on predicting latents, pixels or 3D structure.
    Yann LeCun (chairman), Alexandre LeBrun (CEO), Saining Xie (chief science officer)
  • Omni and natively multimodal models
    One model trained from the start on text, images, audio and video that understands and generates across modalities, often in real time.
    Koray Kavukcuoglu, Qwen team, Muse team
  • Continual learning and memory
    Systems that keep learning after deployment, instead of staying frozen at a training cutoff, through weight updates, test-time training or external memory.
    Ilya Sutskever, Andrej Karpathy, Dwarkesh Patel (framing the bottleneck)
  • Diffusion language models
    Generate blocks of text in parallel by iterative denoising instead of one token at a time, trading some quality for large speed gains.
    Stefano Ermon, Aditya Grover, Volodymyr Kuleshov, Shen Nie, Chongxuan Li and co-authors (LLaDA), Gemini Diffusion and DiffusionGemma teams
  • Efficient architectures: hybrid, linear and sparse attention
    Replace quadratic full attention with linear, recurrent or sparse layers so million-token contexts and long agent loops stay affordable.
    Qwen team, DeepSeek team (NSA, DSA, V4 attention), Kimi Team (Kimi Linear)
  • Automated science and AI scientists
    AI agents that propose hypotheses, run experiments in silico or in autonomous labs, and learn from results to make discoveries with little step-by-step direction.
    AlphaEvolve and co-scientist teams, Tianyi Peng and Claude agents (Fermat formalization), Kevin Buzzard (Lean FLT project)
  • Embodied and robotics foundation models
    General-purpose vision-language-action and world-action models that drive robots across tasks and bodies, improved by deployment experience.
    Sergey Levine, Chelsea Finn, Karol Hausman, Carolina Parada, Helix team
  • Long-horizon autonomous agents
    Agents that work autonomously for hours to days, often in parallel teams, tracked by the length of task they complete rather than benchmark accuracy.
    METR (time-horizon measurement), Nicholas Carlini (agent teams), Dario Amodei
  • Program synthesis, neuro-symbolic systems and ARC-style generalization
    Combine neural intuition with discrete program search or refinement loops to learn new tasks from few examples; ARC Prize tracks progress.
    Francois Chollet, Greg Kamradt, Poetiq team (refinement harnesses)
  • Interpretability-driven alignment
    Read and steer a model's internal circuits and features to verify what it learned, rather than only testing its behavior.
    Dario Amodei, Interpretability team, Eric Ho
  • RL scaling and the era of experience
    Scale RL on verifiable tasks and simulated environments so models learn from trial and error, ultimately from lifelong streams of experience rather than human text.
    David Silver, Richard Sutton, Dario Amodei
  • Automated AI R&D (recursive self-improvement)
    Use AI agents to do AI research itself; OpenAI says it has an 'automated research intern' and targets an automated AI researcher by March 2028.
    Sam Altman, Jakub Pachocki, Dario Amodei, Andrej Karpathy (autoresearch)
  • Small and on-device models
    Capable 1-30B-parameter models that run locally or cheaply, via distillation, MoE and quantization, argued to be the right workhorse for most agent calls.
    Peter Belcak, Pavlo Molchanov, Andrej Karpathy, Gemma team
  • Recursive and latent-space reasoning models
    Tiny networks that loop on a hidden state to solve hard puzzles from ~1,000 examples, an alternative to long chains of thought in huge models.
    Guan Wang and co-authors (HRM), Alexia Jolicoeur-Martineau (TRM)
  • Is scaling over? The 'age of research' debate
    Whether more pretraining and RL compute alone gets to AGI, or whether the bottleneck is now new ideas such as generalization and learning from experience.
    Ilya Sutskever (age of research), Richard Sutton (LLMs a dead end), Yann LeCun (needs world models)
  • Agent containment, monitoring and loss-of-control research
    Detect and contain autonomous agents that cheat, collude or escape sandboxes, via monitoring, chain-of-thought checks and independent incident investigation.
    Chris Painter, Beth Barnes, Ryan Greenblatt, Alignment and security teams