Omni and natively multimodal models

One model trained from the start on text, images, audio and video that understands and generates across modalities, often in real time.

Bolting vision and audio onto text models loses cross-modal reasoning and adds latency. Training every modality together (early fusion) yields any-to-any input and output, live voice and video interaction, and grounding in physical-world knowledge.

Where it stands. By Oct 2026 Google, Meta, Alibaba, DeepSeek, MiniMax and Moonshot all ship natively multimodal models; true any-to-any generation is still staged (video first at Google).

Evidence

For

  • Qwen3-Omni (2025-09-22): one model keeps text, image, audio and video quality equal to same-size single-modal Qwen models, with open-source SOTA on 32 of 36 audio benchmarks (authors' claim).
  • Amazon Nova 2 Omni (2025-12-02): reasoning model taking text, image, video and speech input and producing text and images (preview).
  • Meta's Muse Spark (2026-04-08) is natively multimodal with visual chain of thought; Gemini Omni (2026-05-19) creates from any input, starting with video.
  • Thinking Machines (2026-05-11) trained interaction models from scratch to take audio, video and text continuously and respond in real time.
  • MiniMax M3 (June 2026) trains multimodally from step zero; Moonshot's Kimi K2.5 (Jan 2026) and DeepSeek V4.1-Flash (2026-09-10) are natively multimodal.
  • OpenAI's GPT-Live (2026-08-03) uses a turnless speech model, and GPT-Live-1 (2026-09-10) brings full-duplex voice conversation to the API (OpenAI descriptions).

Against

  • Generation is not unified yet: Gemini Omni launched with video output only, with image and audio output to follow 'in time' (Google, 2026-05-19).
  • Nova 2 Omni was a preview and emits only text and images, not speech or video.
  • Parity claims are self-reported: Qwen3-Omni's 'no degradation' compares only with same-size Qwen single-modal models.
  • Thinking Machines cites a frontier model card saying interactive, hands-on use helped less than autonomous long-running harnesses, so real-time interaction trades against autonomy (2026-05-11).
  • Inference: headline frontier progress in 2026 (ARC-AGI-3, METR horizons, SWE tasks) is measured on text and agent tasks, so omni gains show mainly in voice, video and interface quality.

Milestones

Who is working on it

The labs with the most milestones here are Google DeepMind (2), OpenAI (2), Amazon (1), Meta (1), Alibaba Qwen (Tongyi) (1) and DeepSeek (1).

Sources

  1. blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/
  2. ai.meta.com/blog/introducing-muse-spark-msl/
  3. arxiv.org/abs/2509.17765
  4. aws.amazon.com/about-aws/whats-new/2025/12/amazon-nova-2-omni-preview/
  5. thinkingmachines.ai/blog/interaction-models/
  6. huggingface.co/MiniMaxAI/MiniMax-M3
  7. kimi.com/blog/kimi-k2-5
  8. api-docs.deepseek.com/news/news260910
  9. openai.com/news/rss.xml
  10. rits.shanghai.nyu.edu/ai/moonshot-ai-releases-kimi-k2-5-with-agent-swarm-and-frontier-visi

This research bet was checked against its sources on 6 October 2026. How we check