World models

Learn a predictive model of how the world responds to actions and plan inside it. Rivals disagree on predicting latents, pixels or 3D structure.

Intelligence needs an internal simulator of the physical world as well as next-token prediction over text. Camps differ on form. LeCun's JEPA predicts abstract latents, DeepMind's Genie and Runway generate interactive video, World Labs builds persistent 3D scenes, and robotics groups fuse video models with action prediction.

Where it stands. Heavily funded and fragmented. JEPA (AMI) is pre-product; Genie is a subscriber prototype; World Labs agreed on 2026-09-28 to join AMD (closing late 2026); robotics is converging on world-action models.

Evidence

For

  • V-JEPA 2 (Meta, 2025-06-11): 1.2B-parameter video model pretrained on 1M+ hours of video; 62 hours of robot data gave zero-shot planning and 65-80% pick-and-place on unseen objects (company-reported).
  • Genie 3 (2025-08-05): real-time interactive worlds at 24 fps and 720p, consistent for minutes; DeepMind tested it as a training ground for its SIMA agent.
  • Veo 3 paper (2025-09-24): one video model solved segmentation, edge detection, maze and other tasks zero-shot, argued as a path to generalist vision models.
  • Dreamer 4 (2025-09-29): first agent to obtain Minecraft diamonds from offline data alone, learning behavior by RL inside a learned world model.
  • Robotics is converging: NVIDIA reports a world-action model (DreamZero) at 1750 Elo on RoboArena in Apr 2026 versus 1622 for pi-0.5.
  • Capital: World Labs raised $1B (2026-02-18); AMI Labs raised a $1.03B seed at $3.5B pre-money (2026-03-09).

Against

  • No world-model system yet rivals LLMs on general reasoning or coding; AMI has no product, and TechCrunch notes commercial applications could take years (2026-03-09).
  • Genie limits: Project Genie sessions capped at 60 seconds, with imperfect physics and prompt-following (Google, 2026-01-29); Genie 3 holds coherence for minutes, not hours.
  • The label is blurry: World Labs' own taxonomy (2026-06-03) classes most video models as 'renderers', not simulators or planners.
  • World-action models cost more: about 9 ZFLOPs for action tuning and 50+ with full video pretraining, with slower inference than VLAs (NVIDIA, 2026-06-15).
  • Inference: LLM agents already build in-context world models; GPT-6 Astra wrote compact symbolic models of unseen ARC-AGI-3 games (ARC Prize, 2026-09-03), weakening the case that a new architecture is required.

Milestones

Related launches

  • GWM Worlds 2 Runway · 3 September 2026
    GWM Worlds 2 streams 720p video at 24 fps with 48 kHz synchronized audio, controlled live by text actions, camera motion and multiple users.
  • Solaris Runway · 31 August 2026
    Solaris is an 'interface world model' that renders interactive software interfaces frame by frame in real time, with no code layer behind them.
  • Lucy 2.5 Decart · 16 July 2026
    Lucy 2.5 improves live video editing at 30 fps with physically aware VFX, object removal, global style changes and stronger temporal consistency.
  • Qwen-AgentWorld (language world models for agents) Alibaba (Qwen) · 22 June 2026
    Language world models (35B-A3B and 397B-A17B) that simulate agent environments across seven domains, trained on 10M+ interaction trajectories.
  • Qwen-Robot Suite (RobotNav, RobotManip, RobotWorld) Alibaba (Qwen) · 15 June 2026
    Three foundation models for physical-world intelligence, covering navigation, manipulation and a robot world model.
  • Oasis 3 Decart · 10 June 2026
    Oasis 3 is Decart's interactive world model for physical AI, starting with autonomous driving, available by API on day one.
  • WebWorld (8B, 14B, 32B) Alibaba (Qwen) · 16 February 2026
    Open web-simulator world model (8B, 14B, 32B) trained on 1M+ open-web interactions; training Qwen3-14B on its trajectories lifts WebArena by 9.2%.
  • Lucy 2.0 Decart · 26 January 2026
    Lucy 2.0 transforms live video at 30 fps and 1080p with near-zero latency, using a pure diffusion model with no depth maps or 3D.
  • World API World Labs · 21 January 2026
    World Labs opens a public API for generating explorable 3D worlds from text, images and video, bringing Marble to developers.
  • GWM-1 (General World Model) Runway · 11 December 2025
    GWM-1 is Runway's first general world model. It makes frame-by-frame, action-conditioned real-time video in three variants, Worlds, Robotics and Avatars.
  • RTFM (Real-Time Frame Model) World Labs · 16 October 2025
    RTFM generates video of an interactive 3D world in real time on a single H100, using posed frames as spatial memory.
  • Matrix-Game 2.0 Skywork AI · 18 August 2025
    Matrix-Game 2.0 is an MIT-licensed real-time interactive world model, generating minute-long scenes at 25 fps from keyboard and mouse input.
  • HunyuanWorld 1.0 Tencent · 29 July 2025
    HunyuanWorld 1.0 generates explorable 3D worlds from text or images via panoramic world proxies, with mesh export.
  • MirageLSD Decart · 17 July 2025
    MirageLSD re-skins a live video stream in real time with near-zero latency, the first live-stream diffusion (LSD) model.
  • Muse (World and Human Action Model, WHAM) Microsoft Research · 19 February 2025
    Muse (WHAM) generates both game visuals and controller actions, published in Nature, with weights and a demonstrator released on Azure AI Foundry.
  • Cosmos World Foundation Model Platform NVIDIA · 7 January 2025
    NVIDIA's Cosmos is an open platform of pretrained world foundation models, video curation pipeline and tokenizers for training physical-AI systems.
  • Genie 2 Google DeepMind · 4 December 2024
    Foundation world model that turns one image into a playable 3D environment, consistent for up to a minute.
  • Oasis Decart · 31 October 2024
    Oasis is a real-time, playable Minecraft-like world generated frame by frame by a transformer from keyboard and mouse input.
  • SIMA Google DeepMind · 13 March 2024
    Instructable agent that follows natural-language commands across nine 3D games using only screen pixels and keyboard and mouse.
  • Genie (generative interactive environments) Google DeepMind · 23 February 2024
    11B-parameter world model trained on unlabeled internet videos that turns an image into a playable 2D environment.
  • V-JEPA Meta · 15 February 2024
    Video JEPA learns by predicting masked spatio-temporal regions in latent space from unlabeled video; frozen encoder reused across tasks.

Every launch in the atlas

Who is working on it

The labs with the most milestones and launches here are World Labs (8), Google DeepMind (7), Decart (5), Runway (4), NVIDIA (3) and Alibaba Qwen (Tongyi) (3).

Sources

  1. techcrunch.com/2026/03/09/yann-lecuns-ami-labs-raises-1-03-billion-to-build-world-models/
  2. latent.space/p/ainews-yann-lecuns-ami-labs-launches
  3. ai.meta.com/blog/v-jepa-2-world-model-benchmarks/
  4. deepmind.google/blog/genie-3-a-new-frontier-for-world-models/
  5. blog.google/innovation-and-ai/models-and-research/google-deepmind/project-genie/
  6. worldlabs.ai/blog/taxonomy-of-world-models
  7. worldlabs.ai/blog/amd-announcement
  8. developer.nvidia.com/blog/pretrained-to-imagine-fine-tuned-to-act-the-rise-of-world-action
  9. arcprize.org/blog/astra

This research bet was checked against its sources on 6 October 2026. How we check