Omni and natively multimodal models
One model trained from the start on text, images, audio and video that understands and generates across modalities, often in real time.
Bolting vision and audio onto text models loses cross-modal reasoning and adds latency. Training every modality together (early fusion) yields any-to-any input and output, live voice and video interaction, and grounding in physical-world knowledge.
Where it stands. By Oct 2026 Google, Meta, Alibaba, DeepSeek, MiniMax and Moonshot all ship natively multimodal models; true any-to-any generation is still staged (video first at Google).
Evidence
For
- Qwen3-Omni (2025-09-22): one model keeps text, image, audio and video quality equal to same-size single-modal Qwen models, with open-source SOTA on 32 of 36 audio benchmarks (authors' claim).
- Amazon Nova 2 Omni (2025-12-02): reasoning model taking text, image, video and speech input and producing text and images (preview).
- Meta's Muse Spark (2026-04-08) is natively multimodal with visual chain of thought; Gemini Omni (2026-05-19) creates from any input, starting with video.
- Thinking Machines (2026-05-11) trained interaction models from scratch to take audio, video and text continuously and respond in real time.
- MiniMax M3 (June 2026) trains multimodally from step zero; Moonshot's Kimi K2.5 (Jan 2026) and DeepSeek V4.1-Flash (2026-09-10) are natively multimodal.
- OpenAI's GPT-Live (2026-08-03) uses a turnless speech model, and GPT-Live-1 (2026-09-10) brings full-duplex voice conversation to the API (OpenAI descriptions).
Against
- Generation is not unified yet: Gemini Omni launched with video output only, with image and audio output to follow 'in time' (Google, 2026-05-19).
- Nova 2 Omni was a preview and emits only text and images, not speech or video.
- Parity claims are self-reported: Qwen3-Omni's 'no degradation' compares only with same-size Qwen single-modal models.
- Thinking Machines cites a frontier model card saying interactive, hands-on use helped less than autonomous long-running harnesses, so real-time interaction trades against autonomy (2026-05-11).
- Inference: headline frontier progress in 2026 (ARC-AGI-3, METR horizons, SWE tasks) is measured on text and agent tasks, so omni gains show mainly in voice, video and interface quality.
Milestones
- DeepSeek-V4.1-Flash adds native visual understanding DeepSeek · 10 September 2026
- GPT-Live-1 brings full-duplex voice conversations to the OpenAI API OpenAI · 10 September 2026
- Gemini Omni 1.1 Flash adds scene extension up to 40 s, start/end frames and 4K upscaling Google DeepMind · 27 August 2026
- OpenAI GPT-Live brings continuous, turnless voice interaction with low-latency architecture OpenAI · 3 August 2026
- MiniMax-M3: native multimodal MoE with 1M context via MiniMax Sparse Attention MiniMax · June 2026
- Gemini Omni Flash creates anything from any input, starting with video Google DeepMind · 19 May 2026
- Interaction models research preview for real-time audio, video and text Thinking Machines Lab · 11 May 2026
- Muse Spark is the first natively multimodal reasoning model from Meta Superintelligence Labs Meta · 8 April 2026
- Kimi K2.5: native multimodal model from continued pretraining on ~15T mixed tokens Moonshot AI · January 2026
- Amazon Nova 2 Omni in preview Amazon · 2 December 2025
- Qwen3-Omni technical report describes a single model across text, image, audio and video Alibaba Qwen · 22 September 2025
- Chameleon introduces mixed-modal early-fusion foundation models Meta · 16 May 2024
Who is working on it
- Koray Kavukcuoglu, Google DeepMind (Gemini Omni)
- Qwen team, Alibaba Qwen (Qwen3-Omni)
- Muse team, Meta Superintelligence Labs
- Nova team, Amazon (Nova 2 Omni)
- Interaction models team, Thinking Machines Lab
- M3 team, MiniMax
- V4.1 team, DeepSeek
The labs with the most milestones here are Google DeepMind (2), OpenAI (2), Amazon (1), Meta (1), Alibaba Qwen (Tongyi) (1) and DeepSeek (1).
Sources
- blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/
- ai.meta.com/blog/introducing-muse-spark-msl/
- arxiv.org/abs/2509.17765
- aws.amazon.com/about-aws/whats-new/2025/12/amazon-nova-2-omni-preview/
- thinkingmachines.ai/blog/interaction-models/
- huggingface.co/MiniMaxAI/MiniMax-M3
- kimi.com/blog/kimi-k2-5
- api-docs.deepseek.com/news/news260910
- openai.com/news/rss.xml
- rits.shanghai.nyu.edu/ai/moonshot-ai-releases-kimi-k2-5-with-agent-swarm-and-frontier-visi
This research bet was checked against its sources on 6 October 2026. How we check