Audio, voice and music

Models for speech, voices, sound and music, from transcription and text to speech to music generation. The atlas has logged 134 launches and papers on it since 5 January 2023, 65 of them in 2026.

Google DeepMind, Alibaba Qwen (Tongyi) and OpenAI have the most.

Key launches

The 60 most significant of 134, newest first. Every launch in the atlas

  • SynthID Detector Google · 7 October 2026
    Google opened a public SynthID Detector website that checks images, video and audio for SynthID watermarks from Google, OpenAI, Nvidia and Kakao, with Apple support promised soon.
  • EmbeddingGemma 2 Google DeepMind · 6 October 2026
    Google DeepMind released EmbeddingGemma 2, an open 740M parameter embedding model built on Gemma 4 that maps text, code, images, video and audio into one 768-dimensional space.
  • Griffin Tavus · 1 October 2026
    Tavus unveiled Griffin, a Human Interaction Model for real-time video conversation. In the company's test, 48% of participants thought Griffin-Lite was a real person.
  • Eleven v4 (and v4 Turbo) ElevenLabs · 28 September 2026
    Eleven v4 is ElevenLabs' most emotive TTS: inline delivery tags, consistent multi-speaker dialogue, 90+ languages and a ~150 ms Turbo variant.
  • Kling 4.0 (early access; full launch October 2026) Kuaishou · 28 September 2026
    Kling 4.0: native 30-second clips, up to 10 keyframes, 15 multimodal reference inputs, 10-bit HDR (announced as coming) and a faster Flash tier.
  • UMG and Sony file second suit over Suno v6 Universal, Sony / Suno · 18 September 2026
    Universal and Sony sued Suno again in Massachusetts, covering 60,202 recordings and alleging v6 was trained on outputs of earlier models.
  • Qwen3.8-Omni-Flash Alibaba (Qwen) · 17 September 2026
    Natively multimodal agent model on the Qwen3.8-Next MoE architecture with a 1M-token context, built to plan and finish tasks, not just perceive.
  • UMG and ElevenLabs multi-year licensing agreement Universal Music Group / ElevenLabs · 10 September 2026
    ElevenLabs signed its first major-label deal, a multi-year UMG licence and collaboration that starts with a fan platform for remixes and mashups.
  • GPT-Live 1 OpenAI · 10 September 2026
    GPT-Live 1 is OpenAI's full-duplex voice model. It listens and talks at once while a backend agent reasons and calls tools, at $0.05 per minute.
  • Suno v6 (v6, v6-wild, v6-mini) Suno · 9 September 2026
    Suno v6 is its first model built with the music industry (Warner Music Group, BMG, Believe), with section edits, mashups, sampling and text/audio/image/video prompts.
  • GWM Worlds 2 Runway · 3 September 2026
    GWM Worlds 2 streams 720p video at 24 fps with 48 kHz synchronized audio, controlled live by text actions, camera motion and multiple users.
  • Gemini Omni 1.1 Flash Google DeepMind · 27 August 2026
    Omni 1.1 Flash adds scene extension, first-and-last-frame interpolation, resolution control up to 4K (upscaled) and faster prototyping; GA in the API.
  • Wan3.0 (video) Alibaba (Qwen) · 24 August 2026
    Hosted Wan3.0 makes native 30-second clips up to 1080p with audio and accepts documents, slides, spreadsheets and webpages as references.
  • Wan 3.0 (and Wan 3.0 Prime) Alibaba (Tongyi) · 13 August 2026
    Wan 3.0 generates up to 30 seconds in one pass from text, image, audio, video and documents (PPT, PDF, XLS), at $0.05-$0.20 per second.
  • MiniMax Music 3.0 MiniMax · 13 August 2026
    Music 3.0 composes full songs up to five minutes in one generation, with open weights built from an 8B LLM plus a flow-matching renderer.
  • Suno and BMG global licensing partnership Suno / BMG · 12 August 2026
    Suno signed a global licensing partnership with BMG covering recordings and publishing, settling past use, ahead of its first industry-partnered model.
  • Seedance 2.5 ByteDance · 31 July 2026
    Seedance 2.5 makes 30-second single-pass clips with synced audio, up to 50 multimodal references, region-level edits and 4K, previewed 2026-06-23.
  • MiniMax H3 (Hailuo) MiniMax · 31 July 2026
    MiniMax H3 is an open-weights audio-video model that makes 2K, 15-second clips with native stereo sound, priced below a third of mainstream models at 2K.
  • Munich court rules against Suno in GEMA suit GEMA / Suno · 31 July 2026
    A Munich regional court ruled against Suno in the case brought by German collecting society GEMA.
  • FLUX 3 (early access) Black Forest Labs · 23 July 2026
    FLUX 3: one multimodal flow model for image, video with native audio, and robot-action prediction; open FLUX 3 Dev promised.
  • Qwen-Audio-3.0 and 3.1 (TTS, ASR, realtime duplex) Alibaba (Qwen) · 14 July 2026
    New Qwen-Audio line on Alibaba Cloud with TTS at 200 ms first packet, dialect-aware ASR, and a duplex realtime speech model claimed number one globally.
  • GPT-Live (GPT-Live-1 and mini) OpenAI · 8 July 2026
    GPT-Live is a full-duplex voice model that listens and speaks at once and delegates harder work to a frontier model while it keeps talking.
  • Gemma 4 12B Google DeepMind · 3 June 2026
    Gemma 4 12B: open 12B model with no separate vision or audio encoders, running in 16GB of memory, with native audio input and MTP drafters.
  • Stable Audio 3.0 Stability AI · 20 May 2026
    Stable Audio 3.0: four models (Small SFX, Small, Medium, Large) trained on fully licensed data; Small and Medium are open weights, tracks up to 6:20.
  • Gemini Omni (Omni Flash) Google DeepMind · 19 May 2026
    Gemini Omni Flash makes video from mixed image, audio, video and text input and edits it conversationally; first of a planned family.
  • Gemini Omni Flash Google DeepMind · 19 May 2026
    Gemini Omni Flash takes text, images, audio and video in one prompt and outputs editable video, folding the Veo line into Gemini itself.
  • Flow Music and Lyria 3 Pro Google DeepMind · 19 May 2026
    At I/O 2026 Google Flow Music added Gemini Omni music videos and edits to specific song sections; Lyria 3 Pro, which writes songs up to 3 minutes with intro/verse/chorus…
  • GPT-Realtime-2, GPT-Realtime-Translate and GPT-Realtime-Whisper OpenAI · 7 May 2026
    GPT-Realtime-2 is OpenAI's first voice model with GPT-5-class reasoning, joined by live translation and streaming transcription models.
  • ElevenMusic app ElevenLabs · 29 April 2026
    ElevenMusic is a consumer app for discovering, remixing and creating music from a fully licensed model, launched with 4,000+ independent artists.
  • HappyHorse-1.0 Alibaba · 10 April 2026
    Alibaba revealed HappyHorse-1.0, a joint audio-video model that had topped the Artificial Analysis video arena under an anonymous name.
  • Wan2.7 (T2V, I2V, R2V, video edit) Alibaba (Qwen) · 3 April 2026
    Wan2.7 video suite on Alibaba Cloud covers text-, image- and reference-to-video plus instruction-based video editing; 15-second clips.
  • Qwen3.5-Omni (Plus, Flash and Realtime) Alibaba (Qwen) · 26 March 2026
    Proprietary omni model on the Qwen3.5 generation, with hundreds of billions of parameters, 256K context and claimed leading results on 215 audio and audio-visual subtasks.
  • Suno v5.5 Suno · 26 March 2026
    Suno v5.5 adds Voices (use your own singing), Custom Models trained on your own catalogue (up to three) and a personalisation feature, My Taste.
  • Voxtral TTS Mistral AI · 23 March 2026
    Voxtral TTS is a 4B open-weights text-to-speech model with 3-second voice cloning, 9 languages and about 70ms model latency.
  • MiMo-V2-Pro, V2-Omni and V2-TTS Xiaomi (MiMo) · 18 March 2026
    Xiaomi's first trillion-class model, MiMo-V2-Pro (about 1T total, 42B active), shipped closed alongside an omni model and a speech model.
  • Seedance 2.0 ByteDance · 12 February 2026
    Seedance 2.0 jointly generates audio and video from text, image, audio and video references (9+3+3 per prompt), and Hollywood objected to it.
  • Kling 3.0 Kuaishou · February 2026
    Kling 3.0 unifies video, image, audio and editing, with multi-shot storytelling, native audio, multilingual dialogue and up to 15 seconds per generation.
  • Qwen3-TTS (0.6B, 1.7B) with 12Hz tokenizer Alibaba (Qwen) · 22 January 2026
    Open Apache 2.0 text-to-speech models with 3-second voice cloning and voice design in 10 languages; first-packet latency as low as 97 ms.
  • Wan2.6 series (T2V, I2V, R2V, T2I, image) Alibaba (Qwen) · 16 December 2025
    Wan2.6 adds reference-to-video that keeps a person's look and voice, multi-shot storytelling and clips up to 15 seconds with synchronized audio.
  • Kling 2.6 (native audio, Motion Control) Kuaishou · December 2025
    Kling 2.6 added synchronized audio-video generation and action (motion-control) features, answering Veo 3 and Sora 2 audio.
  • WMG and Suno settle; licensed models promised for 2026 Warner Music Group / Suno · 25 November 2025
    Suno settled with Warner Music Group and agreed to launch licensed models in 2026; downloads move behind paid tiers, and Suno acquired Songkick from WMG.
  • WMG and Udio settle and license Udio's next-generation service Warner Music Group / Udio · 19 November 2025
    Warner Music Group and Udio resolved their litigation with a licensing arrangement for a new Udio service built on artists who opt in.
  • ERNIE 5.0 Baidu · 13 November 2025
    Natively autoregressive omni-modal model trained from scratch on text, image, video and audio with a 2.4T-parameter ultra-sparse MoE (about 3% active).
  • Stability AI alliances with UMG (2025-10-30) and WMG (2025-11-19) Stability AI · 30 October 2025
    Stability AI signed strategic alliances with Universal and Warner Music Group to co-develop licensed, artist-centred professional music tools.
  • UMG and Udio settle, plan licensed 'walled garden' platform Universal Music Group / Udio · 29 October 2025
    Universal Music Group and Udio settled their copyright suit with a payment and licences; Udio disabled downloads ahead of a licensed 2026 platform.
  • LongCat-Flash-Omni Meituan (LongCat) · 23 October 2025
    560B-total (27B active) open omni-modal model with 128K context for real-time audio-visual interaction and streaming speech.
  • Veo 3.1 and Veo 3.1 Fast Google DeepMind · 15 October 2025
    Veo 3.1 adds video extension, first-and-last-frame interpolation and multi-image reference control with audio in every mode; API preview plus Flow.
  • Firefly Image 4 (2025-04) and Image Model 5 (MAX, 2025-10) Adobe · October 2025
    Adobe shipped Firefly Image 4/4 Ultra in April 2025 and Image Model 5, Generate Soundtrack, Generate Speech and a video editor at MAX in October.
  • Sora 2 and the Sora app OpenAI · 30 September 2025
    Sora 2 generates video with synchronized dialogue and sound, better physics, and "cameos" of real people, launched with a social Sora app.
  • Suno Studio Suno · 25 September 2025
    Suno Studio is a browser 'generative audio workstation' with a multitrack timeline, stem generation, BPM/pitch control and MIDI/audio export.
  • Wan2.5-Preview Alibaba (Qwen) · 24 September 2025
    First Wan model with native synchronized audio (dialogue, music, effects) in one pass, up to 10-second 1080p clips; hosted only, no weights.
  • Suno v5 Suno · 23 September 2025
    Suno v5 delivers more natural vocals and a new composition architecture, described as 'best to date', for Pro and Premier first.
  • Qwen3-Omni-30B-A3B Alibaba (Qwen) · 22 September 2025
    Open natively omni-modal 30B-A3B Thinker-Talker MoE: audio latency as low as 211 ms, 19 speech-input and 10 speech-output languages, Apache 2.0.
  • Eleven v3 (alpha, GA 2026-02-02) ElevenLabs · 3 June 2025
    Eleven v3 adds audio tags ([whispers], [laughs]), multi-speaker dialogue and 70+ languages; the most expressive ElevenLabs TTS at launch.
  • Veo 3 Google DeepMind · 20 May 2025
    Veo 3 generates video with synchronized dialogue, sound effects and ambient audio; Google calls it the first video model with native audio.
  • NotebookLM Audio Overviews Google · 11 September 2024
    NotebookLM turns uploaded documents into a two-host podcast-style conversation, powered by Gemini 1.5.
  • Record labels sue Suno and Udio Universal, Sony, Warner (via RIAA) · June 2024
    Major labels, coordinated by the RIAA, sued Suno (Massachusetts) and Udio (New York) for training on copyrighted recordings, seeking up to $150,000 per work.
  • GPT-4o OpenAI · 13 May 2024
    One network handles text, audio and images end to end; GPT-4o costs half of GPT-4 Turbo and reaches free ChatGPT users.
  • Suno v3 Suno · 21 March 2024
    Suno v3 is the first model Suno called 'radio-quality'. It makes two-minute songs with vocals from a text prompt and is free to all users.
  • VALL-E Microsoft Research · 5 January 2023
    VALL-E treats text-to-speech as language modelling over neural-codec tokens and clones a voice from a 3-second sample, trained on 60K hours.

Most active labs