Embodied and robotics foundation models
General-purpose vision-language-action and world-action models that drive robots across tasks and bodies, improved by deployment experience.
Robotics can follow the LLM recipe of pretraining one model on diverse robot and web data, then adapting it to new tasks and bodies. The 2025-26 additions are RL from real deployments (pi*0.6), whole-body humanoid control (Gemini Robotics 2) and video-pretrained world-action models.
Where it stands. Frontier labs and Physical Intelligence ship ever more general policies on company-reported evidence; independent reliability data and revenue are not found. World-action models are the new rival recipe to VLAs.
Evidence
For
- Gemini Robotics 2 (2026-07-30): one VLA controls full humanoids from feet to fingertips, supports multi-robot teamwork, and adapts to a new robot body in a few hours (company-reported).
- pi*0.6 (2025-11-18) improves from real-world deployment via RL; pi-0.7 (2026-04-16, reported) shows compositional generalization, e.g. an air fryer from two loosely related examples.
- DreamZero, a world-action model on a Wan 2.1 video backbone, reached 1750 Elo on RoboArena (Apr 2026) versus 1622 for pi-0.5 (NVIDIA, 2026-06-15).
- Figure's Helix (2025-02-20) was billed as the first VLA with high-rate whole-upper-body humanoid control and multi-robot collaboration (company claim).
- V-JEPA 2 needed only 62 hours of robot data for zero-shot planning (Meta, 2025-06-11).
Against
- Physical Intelligence is reported pre-revenue with no commercialization timeline despite talks at ~$11B valuation (ai2.work, May 2026).
- Amodei (2026-02-13): robotics changes once models learn new body skills on the job, and diffusion adds another year or two.
- World-action models cost more: ~9 ZFLOPs for action tuning and 50+ with full video pretraining, and they infer slower than VLAs (NVIDIA, 2026-06-15).
- Evidence is mostly company demos and arena Elo; independent long-run success rates are not found.
Milestones
- Gemini Robotics 2, ER 2 and On-Device 2: whole-body humanoid intelligence Google DeepMind · 30 July 2026
- NVIDIA: world-action models emerge as a second recipe beside VLAs NVIDIA · 15 June 2026
- pi-0.7 unveiled with compositional generalization (secondary report) Physical Intelligence · 16 April 2026
- pi*0.6: a VLA that learns from experience (arXiv v1) Physical Intelligence · 18 November 2025
- Gemini Robotics 1.5 and ER 1.5: thinking VLA that learns across embodiments Google DeepMind · 25 September 2025
- NVIDIA Isaac GR00T N1: open humanoid robot foundation model NVIDIA · 18 March 2025
- Figure Helix is a VLA for generalist humanoid control Figure AI · 20 February 2025
Related launches
- Qwen-Drive-1.0 (4B) Alibaba (Qwen) · 2 September 2026
Vision-language foundation model for autonomous driving, built on Qwen3.5-4B, unifying 3D perception with visual question answering. - Model Hardware Standard (research preview) Anthropic · 27 August 2026
The Model Hardware Standard is a shared driver spec letting AI agents operate lab and factory instruments over MCP; research preview opens to select labs. - FLUX 3 (early access) Black Forest Labs · 23 July 2026
FLUX 3: one multimodal flow model for image, video with native audio, and robot-action prediction; open FLUX 3 Dev promised. - Robostral Navigate Mistral AI · 8 July 2026
Robostral Navigate, an 8B model, steers robots with one RGB camera and language instructions and scores 76.6% on unseen R2R-CE environments. - Qwen-Robot Suite (RobotNav, RobotManip, RobotWorld) Alibaba (Qwen) · 15 June 2026
Three foundation models for physical-world intelligence, covering navigation, manipulation and a robot world model. - Oasis 3 Decart · 10 June 2026
Oasis 3 is Decart's interactive world model for physical AI, starting with autonomous driving, available by API on day one. - Qwen-VLA (vision-language-action generalist) Alibaba (Qwen) · 28 May 2026
Unified vision-language-action model on Qwen3.5-4B plus a 1.15B flow-matching action decoder, covering manipulation, navigation and trajectory prediction. - Gemini Robotics-ER 1.6 Google DeepMind · 14 April 2026
Gemini Robotics-ER 1.6 improves pointing, counting, success detection and multi-view spatial reasoning, and adds instrument reading, developed with Boston Dynamics. - GWM-1 (General World Model) Runway · 11 December 2025
GWM-1 is Runway's first general world model. It makes frame-by-frame, action-conditioned real-time video in three variants, Worlds, Robotics and Avatars. - MiMo-Embodied-7B Xiaomi (MiMo) · 19 November 2025
Open 7B vision-language model covering both embodied AI and autonomous driving, said to lead 17 embodied and 12 driving benchmarks. - Gemini Robotics On-Device Google DeepMind · 24 June 2025
Gemini Robotics On-Device is a VLA model optimized to run on the robot itself and the first Gemini Robotics model open to fine-tuning via SDK. - V-JEPA 2 Meta · 11 June 2025
1.2B-param video world model pretrained on 1M+ hours of video, then adapted on only 62 hours of robot data for zero-shot pick-and-place planning. - Gemini Robotics and Gemini Robotics-ER Google DeepMind · 12 March 2025
Gemini Robotics, a vision-language-action model built on Gemini 2.0, plus Gemini Robotics-ER for spatial reasoning; reported more than double prior VLAs on generalization. - Cosmos World Foundation Model Platform NVIDIA · 7 January 2025
NVIDIA's Cosmos is an open platform of pretrained world foundation models, video curation pipeline and tokenizers for training physical-AI systems. - Open X-Embodiment / RT-X Google DeepMind · 3 October 2023
Pooled dataset from 33 labs and 22 robot types; RT-1-X beat single-lab models by 50% on average. - RT-2 Google DeepMind · 28 July 2023
A vision-language-action model that emits robot actions as text tokens from a web-pretrained VLM, and it doubles performance on unseen tasks versus RT-1. - RoboCat Google DeepMind · 20 June 2023
Gato-based self-improving robot agent that learns new tasks from 100-1,000 demos and retrains on its own practice data. - PaLM-E Google · 6 March 2023
PaLM-E feeds images, robot state and text as "multimodal sentences" to one embodied LLM; its 562B version set a state of the art on OK-VQA.
Who is working on it
- Sergey Levine, Chelsea Finn, Karol Hausman, Physical Intelligence
- Carolina Parada, Google DeepMind (Gemini Robotics)
- Helix team, Figure AI
- GR00T and Cosmos teams, NVIDIA
The labs with the most milestones and launches here are Google DeepMind (9), NVIDIA (3), Alibaba Qwen (Tongyi) (3), Physical Intelligence (Pi) (2), Black Forest Labs (1) and Decart (1).
Sources
- deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/
- deepmind.google/blog/gemini-robotics-15-brings-ai-agents-into-the-physical-world/
- arxiv.org/abs/2511.14759
- ai2.work/blog/physical-intelligence-nears-11b-valuation-as-pi-0-7-learns-what-it-was-never
- developer.nvidia.com/blog/pretrained-to-imagine-fine-tuned-to-act-the-rise-of-world-action
- figure.ai/news/helix
- dwarkesh.com/p/dario-amodei-2
- arxiv.org/abs/2506.09985
- nvidianews.nvidia.com/news/nvidia-isaac-gr00t-n1-open-humanoid-robot-foundation-model-simu
This research bet was checked and corrected against its sources on 6 October 2026. How we check