Images and video

Models and products that make, edit or understand images and video. The atlas has logged 197 launches and papers on it since 6 February 2023, 58 of them in 2026.

Google DeepMind, Alibaba Qwen (Tongyi) and OpenAI have the most.

Key launches

The 60 most significant of 197, newest first. Every launch in the atlas

  • SynthID Detector Google · 7 October 2026
    Google opened a public SynthID Detector website that checks images, video and audio for SynthID watermarks from Google, OpenAI, Nvidia and Kakao, with Apple support promised soon.
  • EmbeddingGemma 2 Google DeepMind · 6 October 2026
    Google DeepMind released EmbeddingGemma 2, an open 740M parameter embedding model built on Gemma 4 that maps text, code, images, video and audio into one 768-dimensional space.
  • Nano Banana 2.1 Google DeepMind · 6 October 2026
    Google released Nano Banana 2.1, a Flash-tier image generation and editing model with better mask and ink editing, product recontextualization and factual accuracy, now on…
  • ChatGPT visual ads OpenAI · 5 October 2026
    OpenAI announced a visual ad format in ChatGPT that shows labeled image ads next to image generation results, with US testing starting later this month.
  • FLUX 3 Image Black Forest Labs · 1 October 2026
    FLUX 3 Image adds JSON bounding-box layouts and edits that leave pixels outside the box numerically unchanged, native ~4K output, up to ten references.
  • ChatGPT virtual try-on OpenAI · 1 October 2026
    OpenAI added a Try on button to clothing and accessory listings in ChatGPT, which generates images of the user wearing the item from an uploaded selfie.
  • Griffin Tavus · 1 October 2026
    Tavus unveiled Griffin, a Human Interaction Model for real-time video conversation. In the company's test, 48% of participants thought Griffin-Lite was a real person.
  • Kling 4.0 (early access; full launch October 2026) Kuaishou · 28 September 2026
    Kling 4.0: native 30-second clips, up to 10 keyframes, 15 multimodal reference inputs, 10-bit HDR (announced as coming) and a faster Flash tier.
  • Qwen-Image-2.1 (7B) Alibaba (Qwen) · 20 September 2026
    Open image model with a 7B generation component that unifies text-to-image and editing, with native transparent-image generation and editing.
  • GWM Worlds 2 Runway · 3 September 2026
    GWM Worlds 2 streams 720p video at 24 fps with 48 kHz synchronized audio, controlled live by text actions, camera motion and multiple users.
  • Atlas World Labs · 1 September 2026
    Atlas is a multimodal autoregressive diffusion transformer for generating, reconstructing and simulating space. It generates 1 minute of 1440p camera-controlled video.
  • Gemini Omni 1.1 Flash Google DeepMind · 27 August 2026
    Omni 1.1 Flash adds scene extension, first-and-last-frame interpolation, resolution control up to 4K (upscaled) and faster prototyping; GA in the API.
  • Wan3.0 (video) Alibaba (Qwen) · 24 August 2026
    Hosted Wan3.0 makes native 30-second clips up to 1080p with audio and accepts documents, slides, spreadsheets and webpages as references.
  • Wan 3.0 (and Wan 3.0 Prime) Alibaba (Tongyi) · 13 August 2026
    Wan 3.0 generates up to 30 seconds in one pass from text, image, audio, video and documents (PPT, PDF, XLS), at $0.05-$0.20 per second.
  • Seedance 2.5 ByteDance · 31 July 2026
    Seedance 2.5 makes 30-second single-pass clips with synced audio, up to 50 multimodal references, region-level edits and 4K, previewed 2026-06-23.
  • MiniMax H3 (Hailuo) MiniMax · 31 July 2026
    MiniMax H3 is an open-weights audio-video model that makes 2K, 15-second clips with native stereo sound, priced below a third of mainstream models at 2K.
  • FLUX 3 (early access) Black Forest Labs · 23 July 2026
    FLUX 3: one multimodal flow model for image, video with native audio, and robot-action prediction; open FLUX 3 Dev promised.
  • Qwen-Image-3.0 Alibaba (Qwen) · 20 July 2026
    Third-generation Qwen image model focused on realism and authentic detail; hosted via API as qwen-image-3.0-pro (2026-07-20) and qwen-image-3.0 (2026-08-04).
  • Muse Image (and Muse Video preview) Meta · 7 July 2026
    Meta's first in-house image model behaves as an agent (code, web search, self-critique); Muse Video shown as early preview.
  • Oasis 3 Decart · 10 June 2026
    Oasis 3 is Decart's interactive world model for physical AI, starting with autonomous driving, available by API on day one.
  • Gemini Omni (Omni Flash) Google DeepMind · 19 May 2026
    Gemini Omni Flash makes video from mixed image, audio, video and text input and edits it conversationally; first of a planned family.
  • Gemini Omni Flash Google DeepMind · 19 May 2026
    Gemini Omni Flash takes text, images, audio and video in one prompt and outputs editable video, folding the Veo line into Gemini itself.
  • gpt-image-2 (ChatGPT Images 2.0) OpenAI · 21 April 2026
    gpt-image-2 / ChatGPT Images 2.0: an image model with built-in thinking (web search, multi-image output, self-checks), up to 2K.
  • Seedance 2.0 studio cease-and-desists ByteDance · 16 February 2026
    Disney, Netflix, Paramount, Warner Bros. and Sony sent cease-and-desist letters over Seedance 2.0; the MPA called it 'systemic infringement'.
  • Seedance 2.0 ByteDance · 12 February 2026
    Seedance 2.0 jointly generates audio and video from text, image, audio and video references (9+3+3 per prompt), and Hollywood objected to it.
  • Kling 3.0 Kuaishou · February 2026
    Kling 3.0 unifies video, image, audio and editing, with multi-shot storytelling, native audio, multilingual dialogue and up to 15 seconds per generation.
  • FLUX.2 [pro / flex / dev / klein] Black Forest Labs · 25 November 2025
    FLUX.2 pairs a Mistral-3 24B vision-language model with a rectified-flow transformer, and supports up to 10 reference images and 4MP editing, with a 32B open [dev] model.
  • Nano Banana Pro (Gemini 3 Pro Image) Google DeepMind · 20 November 2025
    Nano Banana Pro, built on Gemini 3 Pro, renders legible multilingual text, outputs up to 4K, blends up to 14 images and keeps 5 people consistent.
  • SAM 3 and SAM 3D Meta · 19 November 2025
    SAM 3 segments and tracks every instance matching a text or exemplar concept; SAM 3D reconstructs objects and human bodies from one image.
  • ERNIE 5.0 Baidu · 13 November 2025
    Natively autoregressive omni-modal model trained from scratch on text, image, video and audio with a 2.4T-parameter ultra-sparse MoE (about 3% active).
  • Marble World Labs · 12 November 2025
    Marble, World Labs' first public product, generates persistent 3D worlds from text, images, video or coarse layouts and exports Gaussian splats, meshes or video.
  • Getty Images v Stability AI (UK High Court judgment) Getty Images / Stability AI · 4 November 2025
    UK High Court rejected Getty's secondary copyright claim, holding that Stable Diffusion's weights are not an 'infringing copy', and made only narrow trade-mark findings for…
  • Sora 2 and the Sora app OpenAI · 30 September 2025
    Sora 2 generates video with synchronized dialogue and sound, better physics, and "cameos" of real people, launched with a social Sora app.
  • Nano Banana (Gemini 2.5 Flash Image) Google DeepMind · 26 August 2025
    Gemini 2.5 Flash Image ("nano-banana") blends several images, keeps characters consistent and edits by natural-language instruction; $0.039 per image via API.
  • Genie 3 Google DeepMind · 5 August 2025
    Genie 3 turns a text prompt into a navigable world rendered in real time at 24 fps and 720p, holding consistency for a few minutes.
  • Qwen-Image (20B MMDiT) Alibaba (Qwen) · 4 August 2025
    20B MMDiT image model with native text rendering for English and Chinese, opened under Apache 2.0 with a technical report.
  • Qwen-Image Alibaba (Qwen) · 4 August 2025
    Qwen-Image is a 20B MMDiT image foundation model, released with code and weights, that claims leading results on complex text rendering, especially Chinese.
  • Wan2.2 (T2V-A14B, I2V-A14B, TI2V-5B) Alibaba (Qwen) · 28 July 2025
    Open video diffusion adopts mixture-of-experts. The A14B experts split denoising by timestep, and a 5B hybrid model gives 720P 24fps on a 4090.
  • V-JEPA 2 Meta · 11 June 2025
    1.2B-param video world model pretrained on 1M+ hours of video, then adapted on only 62 hours of robot data for zero-shot pick-and-place planning.
  • Studios v. Midjourney (copyright suit) Disney, Universal, Warner Bros. Discovery / Midjourney · 11 June 2025
    Disney and Universal sued Midjourney over Darth Vader, Minions, Simpsons-style outputs; Warner Bros. Discovery filed its own suit in September 2025.
  • Seedance 1.0 ByteDance · 10 June 2025
    Seedance 1.0 generates 1080p multi-shot video from text or image with strong motion; a top-ranked model on Artificial Analysis at launch.
  • FLUX.1 Kontext Black Forest Labs · 29 May 2025
    FLUX.1 Kontext is one flow-matching model for in-context image generation and editing from text plus image input, and BFL claims it is up to 8x faster than rivals.
  • Veo 3 Google DeepMind · 20 May 2025
    Veo 3 generates video with synchronized dialogue, sound effects and ambient audio; Google calls it the first video model with native audio.
  • Gen-4 Runway · 31 March 2025
    Gen-4 keeps characters, objects and locations consistent across shots from a single reference image, with no fine-tuning.
  • GPT-4o image generation OpenAI · 25 March 2025
    ChatGPT gets native GPT-4o image generation, replacing DALL-E 3: multi-turn editing and inpainting, Pro users first.
  • Wan2.1 (T2V 1.3B and 14B, I2V 14B) Alibaba (Qwen) · 25 February 2025
    Alibaba open-sources a full video-generation suite under Apache 2.0; the 1.3B text-to-video model needs 8.19 GB VRAM, the 14B tops open rivals on company benchmarks.
  • Wan 2.1 Alibaba (Tongyi) · 25 February 2025
    Wan 2.1 open-sources 14B and 1.3B video models under Apache 2.0; the 1.3B runs in 8.2 GB of VRAM.
  • Cosmos World Foundation Model Platform NVIDIA · 7 January 2025
    NVIDIA's Cosmos is an open platform of pretrained world foundation models, video curation pipeline and tokenizers for training physical-AI systems.
  • Veo 2 Google DeepMind · 16 December 2024
    Veo 2 text-to-video with up to 4K resolution and cinematography control, released with an Imagen 3 update and the Whisk remix tool.
  • Genie 2 Google DeepMind · 4 December 2024
    Foundation world model that turns one image into a playable 3D environment, consistent for up to a minute.
  • HunyuanVideo Tencent · 3 December 2024
    HunyuanVideo, a 13B-parameter open-source video model, matched closed models like Runway Gen-3 and Luma 1.6 in Tencent's 1,533-prompt human eval.
  • FLUX.1 [pro / dev / schnell] Black Forest Labs · 1 August 2024
    Black Forest Labs launches with FLUX.1: three 12B image models (pro API, dev open-weight non-commercial, schnell Apache 2.0) and a $31M seed.
  • Segment Anything Model 2 (SAM 2) Meta · 29 July 2024
    Real-time promptable segmentation for images and video, ~44 fps, with SA-V dataset of 51K videos and 600K+ masklets; Apache 2.0.
  • Gen-3 Alpha Runway · 17 June 2024
    Gen-3 Alpha, trained on new large-scale multimodal infrastructure, markedly improves fidelity, motion and temporal control and human faces.
  • Kling (1.0) Kuaishou · June 2024
    Kuaishou's Kling opened as a beta in China, a Sora-like text/image-to-video model that was publicly usable months before Sora.
  • GPT-4o OpenAI · 13 May 2024
    One network handles text, audio and images end to end; GPT-4o costs half of GPT-4 Turbo and reaches free ChatGPT users.
  • Stable Diffusion 3 (early preview) Stability AI · 22 February 2024
    Stable Diffusion 3 was previewed as a diffusion transformer trained with flow matching, at 800M to 8B parameters and waitlist-only.
  • Sora (technical preview) OpenAI · 15 February 2024
    OpenAI previewed Sora, a diffusion transformer that generates up to one-minute HD video from text and was shown only to red teamers and select creatives.
  • Stable Diffusion XL 1.0 Stability AI · 26 July 2023
    SDXL 1.0 ships with open weights, a 3.5B-parameter base plus refiner and native 1024x1024 output, and runs on 8GB consumer GPUs.
  • ControlNet Stanford University · 10 February 2023
    ControlNet adds spatial control (edges, depth, pose, segmentation) to pretrained text-to-image diffusion models without retraining the base model.

Most active labs