How AI learned to make images and video

What the two parts share

Image generators came first, video generators reuse their parts, and world models such as Genie 3 and V-JEPA 2 are the attempt to make machines understand a scene as well as render it.

Four ideas recur in both halves. A compressed latent space makes generation affordable, a transformer replaces the convolutional network as models scale, text understanding matters more than image quality, and the data behind all of it is what the lawsuits are about. The images half runs from 2014 to October 2026 and the video half from 2016. Read them in order for the arc, or jump to the mechanisms sections for the explanations.

2014 to 2019

GANs, and the first diffusion paper

GANs, which train a generator network against a critic network, led image generation for six years, and diffusion's first paper was ignored for five.

The bar in Montreal

In 2014 Ian Goodfellow, then a PhD student under Yoshua Bengio, was at a Montreal bar called Les 3 Brasseurs when friends discussed generating photos with a heavy statistical method. According to MIT Technology Review, he proposed something simpler. One network would generate images and another would try to tell them from real ones, each improving by exploiting the other's mistakes. He coded it that night and it worked on the first try. The result was the generative adversarial network, published in June 2014. (Jürgen Schmidhuber has long disputed priority; the field's usual answer is that his earlier ideas were related but distinct.)

GANs were fast, because generation is a single forward pass, and sharp. They were also unstable to train, prone to mode collapse (covering only part of the data) and clumsy at following text. StyleGAN (December 2018, Nvidia's Tero Karras and colleagues) showed how far they could go. A mapping network whose "style" modulates every layer gave scale-by-scale control and photorealistic faces, trained on 1024-pixel images for about 41 V100-days.

The paper that was ignored

In March 2015 Jascha Sohl-Dickstein and colleagues published "Deep Unsupervised Learning using Nonequilibrium Thermodynamics," the first diffusion model, which gradually destroys an image with noise and then learns to reverse the process. Quanta Magazine reports that his own samples were so poor that he joked one blob "looks like a truck." Almost nobody built on it for about five years. Also in this period, VQ-VAE (November 2017) turned images into grids of discrete tokens from a learned codebook, which made image generation look like language modelling and set up DALL·E. The deep dive is B11.

2020 to 2021

Diffusion works, and volunteers build art tools around CLIP

Song and Ermon's score matching and Ho, Jain and Abbeel's DDPM made diffusion work, OpenAI released CLIP and withheld the DALL·E model, and volunteers built the first AI art tools from what was public.

Score matching and DDPM

In 2019 Yang Song and Stefano Ermon at Stanford published score matching, a way to learn the direction toward more probable images, without having read Sohl-Dickstein's paper. In June 2020 Jonathan Ho, Ajay Jain and Pieter Abbeel's DDPM tied the two views together and made diffusion practical, reaching 3.17 FID on CIFAR-10. DDIM (October 2020) made sampling 10 to 50 times cheaper, and Song, Sohl-Dickstein and others unified everything as stochastic differential equations in November.

DALL·E 1 and CLIP

On 5 January 2021 OpenAI announced DALL·E and CLIP. DALL·E was a 12-billion-parameter transformer, trained on 250 million text-image pairs with 1,024 V100 GPUs, that predicted image tokens one at a time; samples were reranked by CLIP over 512 candidates. The transformer was never released; only the small tokenizer was. CLIP, trained on 400 million web pairs so that images and their captions land close together, was released. It turned out to matter more, because nearly every image system that followed used it to connect text and images.

Volunteers build AI art tools from CLIP

Between January and mid-2021, before any lab shipped an image generator to the public, a community of volunteers built one from the released parts. Big Sleep, then VQGAN+CLIP (using the sharp tokenizer from the CompVis lab that built Stable Diffusion a year later), then CLIP-guided diffusion, then Disco Diffusion (built on OpenAI's released guided-diffusion weights plus Katherine Crowson's own fine-tune) circulated as free Google Colab notebooks and in Discord servers such as EleutherAI's, where a VQGAN-CLIP bot lived in a channel called the-faraday-cage. Ryan Murdock, Crowson and others wrote the code; the results were abstract and dreamlike, and for much of the public this was the first AI art. Around the same time Boris Dayma's DALL·E mini, built during a Hugging Face and Google JAX event, became the first widely shared open clone (later Craiyon).

Diffusion beats GANs

In May 2021 Prafulla Dhariwal and Alex Nichol's "Diffusion Models Beat GANs on Image Synthesis" introduced classifier guidance and settled the argument on benchmarks. In December Jonathan Ho and Tim Salimans's classifier-free guidance, which trains one network with the prompt randomly blanked and then pushes along the difference at generation time, removed the need for a separate classifier. OpenAI's GLIDE showed humans preferred it. The same month, on the same day, the Latent Diffusion paper from CompVis in Munich arrived, and it is the foundation of the next section.

2022

DALL·E 2, Imagen, Midjourney and Stable Diffusion

DALL·E 2, Imagen, Midjourney and Stable Diffusion all appeared within five months, and Stable Diffusion, the one with open weights, ran on a home gaming card.

DALL·E 2, Imagen and Parti

OpenAI's DALL·E 2 (6 April; wider release in September) was the first image generator of mass-market quality. Google answered with Imagen in May, which found that a frozen text-only language model (T5) understood prompts better than CLIP and that making the text encoder bigger helped more than making the image model bigger, and with Parti in June, a 20-billion-parameter autoregressive model. Google released neither. Imagen's alumni, Ho, William Chan, Chitwan Saharia and Mohammad Norouzi, left to found Ideogram.

Midjourney

David Holz, formerly of the hand-tracking company Leap Motion, opened Midjourney's beta on Discord on 12 July. The Register reported a month later that it was self-funded, profitable, about ten people, and deliberately biased away from photorealism to avoid deception. In August Jason Allen's Midjourney image won a Colorado State Fair art prize, setting off the first major backlash from artists.

Stable Diffusion

On 22 August 2022 Stability AI released Stable Diffusion v1.4 weights, under a permissive licence, built by the CompVis group (Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Björn Ommer) with Runway and LAION. It ran in about 7 GB of graphics memory, so a gaming card could do it. Latent diffusion runs the denoising in a compressed space. A 512×512×3 image (about 786,000 numbers) becomes a 64×64×4 latent (about 16,000), which is why it was cheap enough to run at home.

Within weeks the community had added DreamBooth (teach the model your face or style from a few photos) and textual inversion. Stability raised $101 million in October. Runway then posted Stable Diffusion 1.5 and Stability filed a takedown, calling it a leak, before backing down; Runway's CEO cited Stability's compute donation. The LAION dataset behind it, 5.85 billion image-text pairs published in March, was open, and open data raised legal questions that are still being argued.

Compute figures that don't agree

Stability's launch post described a 4,000-A100 cluster; the model card says about 150,000 A100-hours. Wikipedia reports roughly $600,000 for training; PixArt's authors later claimed their own model needed 675 A100-days against 6,250 for SD 1.5. Whichever is right, a model that cost under a million dollars to train nearly sank its sponsor. The Register, citing Forbes, reported Stability projecting about $11 million of 2023 revenue against about $99 million a year in cloud costs. The deep dive is B11.

2023 to 2024

LoRA, ControlNet, DALL·E 3 captions and the move to transformers

Open-source developers built LoRA and ControlNet, OpenAI showed with DALL·E 3 that better captions mattered more than a better architecture, and the U-Net gave way to the transformer in Stable Diffusion 3 and FLUX.1.

LoRA, ControlNet and the open toolkit

LoRA, ported to diffusion by Simo Ryu in December 2022, shrank personalization from multi-gigabyte checkpoints to adapters of 1 to 6 MB. ControlNet (Lvmin Zhang and Maneesh Agrawala, February 2023) let you steer composition with edges, depth maps or poses and became the open community's killer app. Tools like AUTOMATIC1111's web interface and ComfyUI, and the Civitai model-sharing site (which took Andreessen Horowitz money in November 2023), grew into an ecosystem; Civitai also became known for deepfake requests, and NovelAI's leaked code fed anime-style models.

SDXL and the DALL·E 3 captions

Stable Diffusion XL (July 2023) was the last big U-Net generation, three times larger with two text encoders. DALL·E 3 (October 2023, inside ChatGPT) gained its prompt-following from better captions. OpenAI rebuilt its training captions with a trained captioner (about 95% synthetic), so the model learned from accurate descriptions and followed prompts far better. SD3 later used a 50/50 mix of real and synthetic captions.

Real-time diffusion and the LAION-5B findings

Latent consistency models and SDXL Turbo (October and November 2023) cut generation from tens of steps to one to four, so images appeared as you typed. In December 2023 the Stanford Internet Observatory found links to child sexual abuse material in LAION-5B, and LAION pulled the dataset; Re-LAION-5B returned in August 2024 with 2,236 links removed. In August 2024 Hugging Face removed Stable Diffusion 1.5, which IEEE Spectrum linked to the finding.

DiT and Stable Diffusion 3

In December 2022 William Peebles and Saining Xie's DiT showed a plain transformer over patches of the latent could replace the U-Net, with quality that improved smoothly with compute. Stable Diffusion 3's paper (March 2024) combined it with rectified flow, which trains the network to follow near-straight paths from noise to image so fewer steps are needed, and with a design (MMDiT) giving text and image tokens their own weights but joint attention. In March 2024 Emad Mostaque resigned as Stability's CEO.

Black Forest Labs and FLUX

On 1 August 2024 Rombach, Esser and Blattmann, the people behind latent diffusion and Stable Diffusion, launched Black Forest Labs with a $31 million seed and the 12-billion-parameter FLUX.1, a flow transformer with open-weights variants. FLUX.1 [dev] was guidance-distilled and [schnell] used adversarial distillation to run in a few steps. The creators of Stable Diffusion had left the company that sponsored it, and their new model beat it.

VAR brings back autoregression

VAR, from ByteDance and Peking University (April 2024), predicted a whole coarse-to-fine image scale per step instead of a token at a time and cut ImageNet FID from 18.65 to 1.73 with scaling laws like a language model's. It won a NeurIPS 2024 best-paper award. A lawsuit ran alongside it. In mid-2024 ByteDance sued a former intern for 8 million yuan over alleged sabotage of its training infrastructure (the outcome is unverified). In February 2024 Google paused Gemini's people-image generation after over-tuned diversity prompts produced historically inaccurate images.

2025

GPT-4o and Gemini generate images natively

GPT-4o and Gemini generate images inside the same model that reads and writes text.

GPT-4o and the Studio Ghibli portraits

On 25 March 2025 OpenAI turned on native image generation in GPT-4o, after showing it ten months earlier. Per its system card it is autoregressive and embedded in the chat model, unlike diffusion DALL·E. It rendered text and followed long instructions far better, and it allowed studio styles while refusing living artists' styles, which is why Studio Ghibli-style portraits flooded the internet within days. OpenAI reported 130 million users making 700 million images in the first week (one secondary source; treat with caution).

Nano Banana

Google had shipped experimental native image output in Gemini 2.0 Flash in March. On 26 August Gemini 2.5 Flash Image, known as "Nano Banana," could keep a person's face while changing the scene, and consistent editing became the new frontier. It appeared anonymously on the LMArena leaderboard from 12 August, and its nickname reportedly came from a Google product manager's nicknames. Each image cost 1,290 output tokens, about four cents. Nano Banana Pro (20 November, built on Gemini 3 Pro) added 4K output, legible multilingual text, up to 14 input images and search grounding.

Qwen-Image, Seedream, HunyuanImage and FLUX.2

Alibaba's Qwen-Image (August; 20 billion parameters, Apache 2.0) was strong at rendering Chinese and English text. ByteDance's Seedream 4.0 (September) reported adversarial distillation plus speculative decoding for about 1.8 seconds per 2K image. Tencent's HunyuanImage 3.0 (September) is an open 80-billion-parameter mixture-of-experts model that both understands and generates images. Black Forest Labs' FLUX.2 (25 November) paired a 32-billion-parameter flow transformer with a 24-billion-parameter Mistral vision-language model as its text encoder, and the company raised $300 million at a $3.25 billion valuation. The deep dive is B11.

Disney sues Midjourney, and the UK court rejects most of Getty's case

On 11 June 2025 Disney and Universal sued Midjourney, and Warner Bros. followed in September. On 4 November the UK High Court largely rejected Getty Images' case against Stability. The judge held that model weights are not an "infringing copy". Getty had dropped its training claim because training occurred outside the UK, and its trademark wins were narrow, limited to watermarks in certain versions. Getty was given permission to appeal.

2026

OpenAI's models take the top three image rankings

By October 2026 the three highest-ranked image models on the arena.ai leaderboard are OpenAI's, and Artificial Analysis ranks the best open-weights model eighteenth.

Image rankings in October 2026

OpenAI's gpt-image-2 (21 April) added an optional reasoning step and 2K output; its Images 2.5 (8 September) split into Flare (fast) and Sunburst (precise). On the arena.ai leaderboard on 5 October, ranks 1 to 3 were gpt-image-2.5 Sunburst, Flare and gpt-image-2, followed by xAI's Grok Imagine 2.0, Microsoft's MAI-Image, Reve, Meta's Muse Image, Google's Nano Banana 2 and Pro, Seedream 5.0 and Qwen-Image-3.0-Pro. Artificial Analysis shows the same top three and ranks the best open-weights model, Qwen-Image-2.1, eighteenth. Neither top-15 list includes Midjourney or FLUX. Google's Nano Banana 2 (26 February) became the default in Gemini, Flow and Search.

Midjourney V8, FLUX 3, Stability and Ideogram

Midjourney V8 (17 March) was about five times faster, with native 2K output at a higher cost and better quoted text. Black Forest Labs announced FLUX 3 on 23 July as one flow model for images, video with audio and robot actions, with the image model shipping on 1 October; its 7-billion-parameter FLUX 3 Action model is open. Stability pivoted toward audio, video and enterprise. Ideogram released open-weights models in June.

Courts and regulators in 2026

The Grok "undressing" scandal led Ofcom to open an Online Safety Act investigation of X in January. The US Supreme Court declined in March to hear Thaler's challenge to the requirement of human authorship for copyright. Andersen's artists' case against Stability and Midjourney has a summary-judgment hearing now set for February 2027, and Disney's case against Midjourney is in discovery. The C2PA content-credentials standard now has Google, OpenAI, Meta and Amazon aboard.

How image generation works

A current image model is built from a compressed latent space, a text encoder, a denoising network and a method for cutting the number of steps, and each is described below with what it costs.

From GANs to diffusion

  • GANs pit a generator against a discriminator. They are fast and sharp, unstable to train and poor at text.
  • VQ tokenizers compress an image to a grid of integer codes from a learned codebook, so a transformer can generate images as it generates words. DALL·E 1's tokenizer used 32×32 tokens from an 8,192-entry codebook.
  • Diffusion adds noise over many steps, then trains one network to predict the noise at any level; generation starts from static and denoises repeatedly. Training is plain regression, so it is stable and covers the whole data distribution. The price is many sequential calls.
  • Latent diffusion does the denoising in a compressed space, which is why consumer GPUs could run it. The compression caps fine detail such as small text, hence later models' larger latents.

Telling the model what to draw

  • CLIP versus T5. CLIP's text side is cheap and visual but weak at counting, spatial relations and spelling. Imagen found a text-only language model understood prompts better. SD3 combines both, FLUX.2 uses a 24-billion-parameter vision-language model and Qwen-Image uses Qwen2.5-VL.
  • Guidance extrapolates from the model's unconditional guess toward the prompt-conditioned one, trading diversity for adherence at twice the compute. FLUX.1 [dev] bakes guidance into its weights.

Faster and bigger

  • U-Net versus DiT. A U-Net is a convolutional encoder-decoder with good image priors but awkward scaling. DiT cuts the latent into patches and runs a plain transformer, so it inherits language-model tooling and scales smoothly. FLUX.1 is 12 billion parameters, Qwen-Image 20 billion and FLUX.2 32 billion.
  • Flow matching trains the network to move from noise to data along near-straight paths, so sampling needs fewer steps.
  • Few-step distillation trains a student to match a many-step teacher in one to four steps; adversarial versions keep the output sharp, so the GAN critic returns as a patch on diffusion.
  • Autoregressive and native generation emits image tokens inside the language model itself. How the final pixels are decoded in GPT-4o and Gemini is undisclosed.

Lesser-known stories about images

Eleven backstories from the research notes, from the truck-shaped blob in the first diffusion paper to Disney's $1 billion and the end of Sora.

  • Diffusion's founding blob. The paper that started diffusion had samples so poor that its author joked one looked like a truck.
  • The volunteers got there first. Free Colab notebooks and a Discord channel produced the first AI art years before the products.
  • A teacher started LAION. Christoph Schuhmann, a German educator, began it on Discord after DALL·E and reached 400 million pairs in about three months, funded partly by about $9,000 to $10,000 from Emad Mostaque, who later trained Stable Diffusion on a subset.
  • "Data laundering." Andy Baio argued in 2022 that Stability funded the university compute and the dataset while a research-only framing sheltered commercial use.
  • Greg Rutkowski appeared 15 times in a 12-million-image sample, despite dominating prompts; Pinterest was 8.5% of the sample.
  • A real artist was cloned. A Reddit user trained a DreamBooth model on 32 images of illustrator Hollie Mengert and published it under her name.
  • The SD 1.5 fight nearly made Runway and Stability enemies, and a Runway researcher later helped found a company with Stability's creators.
  • Midjourney is nearly all Discord and few people. Its revenue was reported at about $200 million by 2023, and Meta licensed its technology in August 2025.
  • Captions beat architecture. DALL·E 3's jump came from re-captioning the training set, which is why OpenAI would not release the details.
  • A research prize had a sabotage lawsuit attached. VAR won best paper at NeurIPS 2024 while its first author's former employer was suing an intern.
  • Sora's end and Disney. Disney had committed $1 billion and reportedly learned of Sora's shutdown less than an hour beforehand.

Open questions in image generation

Six questions the sources could not settle, from how frontier image models are built to who pays for inference.

  • Architecture opacity. OpenAI and Google have not said whether frontier image generation is autoregressive plus a diffusion decoder or something else, for gpt-image-2 or for Gemini's image output. It is also unknown whether reasoning before drawing helps beyond text and layout.
  • Who sustains the specialists. Midjourney is bootstrapped and Black Forest Labs is moving toward video and robotics. The top-15 lists contain no open-weights model.
  • Copyright. The UK High Court held that weights are not copies (appeal pending), and Andersen's theory that the model itself is a copy survived pleading. Human authorship is still required for copyright.
  • Safety and provenance. Open weights and fine-tune marketplaces sit against platform liability, and watermarks are voluntary and easy to strip.
  • Evaluation. FID assumes Gaussian features and is considered unreliable for text-to-image. CLIPScore is weak on composition and PickScore predicts human preference better. Arenas rank blind votes, but they are now funded companies and have faced gaming criticism.
  • Compute versus value. Who pays for inference is unsettled, since the first week of GPT-4o images overheated GPUs.

2016 to 2022 · Video GANs, Make-A-Video and Imagen Video

Video generation was a curiosity for six years. Then diffusion arrived, and Meta's Make-A-Video and Google's Imagen Video and Phenaki were impressive results that neither company shipped.

The long pre-history

The first deep video generator that got noticed was a 2016 video GAN from Carl Vondrick, Hamed Pirsiavash and Antonio Torralba. It made roughly one-second clips of beaches and train stations, with the foreground and background generated by separate streams. In 2020 NeRF showed that a whole 3D scene could be stored inside a small neural network, the seed of the persistent-world line that returns in 2025. In 2021 VideoGPT, by Wilson Yan, Yunzhi Zhang, Pieter Abbeel and Aravind Srinivas, compressed video into discrete codes and trained a GPT on them, the "tokenize, then predict" template that later systems kept. (Srinivas later ran Perplexity.)

Diffusion enters video

In April 2022 Jonathan Ho and colleagues published Video Diffusion Models, stretching the image diffusion network into time and training on images and video together. Images teach appearance and video teaches motion. A month later Tsinghua and BAAI released CogVideo, a 9-billion-parameter model billed as the first large open text-to-video model, and the first step in what has become a Chinese lead in open video models.

On 29 September 2022 Meta announced Make-A-Video, which solved the missing-dataset problem by learning appearance from images and motion from unlabeled video, with no paired text-and-video data. On 5 October Google published two papers on the same day, Imagen Video (a cascade of models, each sharpening the last) and Phenaki (long, story-like clips from a changing prompt). Both were papers and demos only. Google would not ship a product until Veo in May 2024, which fits a recurring Google pattern of publishing first and shipping late. Four of Imagen Video's authors, Ho, William Chan, Chitwan Saharia and Mohammad Norouzi, later co-founded Ideogram.

U-ViT and DiT

Two papers settled what video models would be made of. Tsinghua's U-ViT (September 2022) and then DiT, the Diffusion Transformer by William Peebles and Saining Xie (December 2022), showed that a plain transformer over patches of the image could replace the convolutional U-Net that had defined diffusion. DiT's quality improved predictably with compute. It is the block underneath Sora and most of what followed. (U-ViT had the idea three months earlier, and Tsinghua's role is under-told; a story that DiT was rejected at a conference first is unverified.) The deep dive is B13.

2023

Runway and Pika launch products, and the Will Smith spaghetti test begins

Runway shipped Gen-2 and Pika launched Pika 1.0, a grotesque clip of Will Smith eating spaghetti became the field's informal test, and Stability released open video weights.

Runway and the first tools

Runway, founded in 2018 by NYU graduate students Cristóbal Valenzuela, Anastasis Germanidis and Alejandro Matamala (a Chilean, a Greek and a Chilean who met at NYU's interactive-media programme), had already helped make Stable Diffusion. It co-wrote the latent-diffusion work and shipped version 1.5 in October 2022. In February 2023 it published Gen-1, which transformed existing video, and shipped Gen-2, text-to-video, by June. Patrick Esser, who led the Gen-1 work, later co-founded Black Forest Labs with Robin Rombach and Andreas Blattmann.

In late November Pika, started by two Stanford AI PhD students, Demi Guo and Chenlin Meng, launched Pika 1.0 with $55 million. Stability released Stable Video Diffusion with a paper whose main contribution was curation, filtering 577 million clips (about 212 years of video) down to 152 million.

The spaghetti test

In March 2023 a clip made with the open ModelScope model, showing a grotesque Will Smith eating spaghetti, went viral on a Stable Diffusion forum. It became a running benchmark for the field, because you could watch each new model do the same prompt better. In February 2024 Smith posted his own parody. A meme is a poor measurement, but progress in this area is often judged in public, by eye.

Gaussian splatting and W.A.L.T

In August 2023 3D Gaussian Splatting replaced NeRF's neural network with millions of explicit, semi-transparent 3D blobs that a graphics card can draw directly, at 30 frames per second or more at 1080p. Because a scene made of splats is a file that can be saved, edited and exported, it became the format persistent worlds would use, unlike frame-by-frame video, which forgets what it just showed. In December 2023 Google and Stanford's W.A.L.T, with Fei-Fei Li among its authors, applied a transformer to video latents and became the nearest public ancestor of Sora.

2024

OpenAI previews Sora while Kling, Vidu, Hailuo and Luma ship products

OpenAI previewed Sora on 15 February as a world simulator and released it on 9 December, ten months later, while Kling, Vidu, Hailuo and Luma shipped products in between.

Sora, February 2024

On 15 February 2024 OpenAI previewed Sora with a report titled "Video generation models as world simulators." The report's recipe was to cut compressed video into spacetime patches (small 3D blocks flattened into tokens), train one model on any length, resolution and aspect ratio, and scale it. The clips, up to a minute long, looked years ahead of anything public. (OpenAI's pages were not retrievable for this chapter; technical details here come from secondary sources and are flagged accordingly.)

Four days later Yann LeCun called pixel-level world modelling "doomed to failure" and Sora a dead end, and the argument is still running. The same week Google DeepMind's Genie (11 billion parameters) learned to make playable 2D worlds from unlabeled internet video, discovering "latent actions" without labels.

The data question and Air Head

In March Mira Murati told the Wall Street Journal that Sora was trained on publicly available and licensed data, and was unsure whether YouTube, Instagram or Facebook videos were included. The same month the Toronto studio shy kids released "Air Head," a short film about a man with a balloon for a head. The fxguide analysis of how it was made found about 300 minutes of generated footage for every minute used, no persistent character, camera moves obeyed about six times in ten, and heavy post-production, including rotoscoping balloon faces and upscaling from 480p.

Kling, Vidu and Hailuo ship while Sora waits

While Sora stayed in preview, Chinese labs shipped. Vidu (Shengshu and Tsinghua, built on U-ViT) arrived in April. Kling from Kuaishou followed in June. It was a DiT with its own 3D video autoencoder and the first public rival at Sora's level, though early access required a Chinese phone number, and fake Kling sites later spread malware. MiniMax's Hailuo came in September. Luma's Dream Machine opened to everyone on 12 June and Runway Gen-3 Alpha followed days later, so the claim that Kling was the first usable Sora-class model is contestable. Google showed Veo at I/O in May.

Scraping leaks, GameNGen and Movie Gen

On 25 July 404 Media published an internal Runway spreadsheet listing more than 3,900 YouTube channels to scrape, with the Gen-3 project codenamed Jupiter. Runway had publicly cited only curated internal datasets. In August, leaked messages reported by the same outlet described Nvidia harvesting YouTube and Netflix video for its Cosmos project; Nvidia said it respected copyright.

August also brought GameNGen, a diffusion model trained on an AI agent playing Doom that simulated the game at 20 frames per second, the first sign that a video model could be interactive. In October Oasis (Decart and Etched) did the same for Minecraft in real time, and Meta announced Movie Gen, a 30-billion-parameter video model with a 13-billion-parameter audio model, as research only. Tim Brooks, a Sora co-lead since January 2023, left OpenAI for Google DeepMind on 3 October, saying he would build world simulators there.

Sora launches to paying users, and beta artists leak it

Sora reached paying users on 9 December. A month earlier a group of its beta artists had leaked access on Hugging Face to protest what they called "art washing"; it was revoked within hours. Google answered in December with Veo 2 and Genie 2, and Tencent released the 13-billion-parameter HunyuanVideo with open weights. The deep dive is B13.

2025

Veo 3 adds sound, Wan runs on gaming cards, and Sora 2 launches a feed

Google's Veo 3 generated sound together with the video, Alibaba's Wan 2.1 ran on consumer graphics cards, and OpenAI's Sora 2 launched as a social video app.

Open weights and Wan

Nvidia released the open Cosmos world models for robotics and cars in January 2025. In February Alibaba open-sourced Wan 2.1 under Apache 2.0, and its 1.3-billion-parameter model ran in about 8 GB of video memory, so a gaming laptop could generate video. ComfyUI added support within two days. Wan 2.2 (July) used a mixture-of-experts design. The public repository shows only versions 2.1 and 2.2, so later Wan versions appear to be closed; that is an inference from the repository, and Alibaba has not confirmed it.

Veo 3 and native audio

On 20 May Google released Veo 3, the first major model to generate dialogue, sound effects and ambient audio together with the video. Native audio became standard within months (Sora 2, Kling, Seedance 2.0, the open LTX-2). ByteDance's Seedance 1.0 followed on 10 June with multi-shot clips, and Midjourney's first video model on 18 June, a week after Disney and Universal sued it for copyright infringement.

V-JEPA 2 and Genie 3

On 11 June Meta released V-JEPA 2, LeCun's answer to Sora. It predicts in an abstract representation instead of pixels, was pretrained on over a million hours of video, and was then adapted with under 62 hours of robot data to plan a robot arm's moves with no task-specific training. Whether that is better than generating pixels is the argument taken up in the world-simulator section below. In August Google DeepMind's Genie 3 generated navigable worlds at 720p and 24 frames per second in real time, with about a minute of visual memory.

Sora 2, Vibes and the MiniMax lawsuit

On 30 September OpenAI launched Sora 2 and a TikTok-style app with "cameos," verified likenesses of real people. Meta had launched Vibes, its own feed of AI videos, five days earlier. Sora 2 shipped with an opt-out copyright policy but an opt-in likeness policy; three days later Sam Altman promised rights-holders more control and revenue sharing. Watermark removers appeared within a week. In September Disney, NBCUniversal and Warner Bros. Discovery sued MiniMax over its Hailuo video app, the first studio suit aimed at a video generator.

Runway Gen-4.5 and the Disney deal

Runway's Gen-4.5 took the top of the Artificial Analysis blind-vote ranking on 1 December, ahead of Google's Veo 3 and OpenAI's Sora 2 Pro. On 11 December Disney announced a $1 billion commitment to OpenAI and a licence for more than 200 characters on Sora. The deep dive is B13.

2026

OpenAI ends Sora, studios object to Seedance 2.0, and money goes to world models

OpenAI announced the end of Sora on 24 March, studios sent cease-and-desist letters over Seedance 2.0, and money went to interactive and 3D world models.

OpenAI ends Sora

On 24 March 2026 OpenAI announced the end of Sora; the app closed on 26 April and the API was scheduled to retire on 24 September. OpenAI gave no explicit reason beyond redirecting the team toward world-simulation research for robotics. Reporting attributes it to economics, with about $1 million a day in running costs, a user base that peaked near a million and fell below half that, and a chip shortage. Disney's $1 billion reportedly never moved, and Disney reportedly learned of the shutdown abruptly (terms were not public).

Seedance 2.0 and the studios

ByteDance's Seedance 2.0 (12 February) generated audio and video jointly from many reference images, clips and sounds. Within 72 hours a clip of Tom Cruise fighting Brad Pitt made by the director Ruairi Robinson had 1.8 million views, the screenwriter Rhett Reese said one person could soon make a studio-grade film, Disney sent a cease-and-desist letter, and Paramount, Netflix, Warner Bros. and Sony followed over the next ten days. No Seedance lawsuit had been reported by mid-July 2026. Seedance 2.5 was previewed in June (30-second clips, up to 50 references, per ByteDance; details unverified).

Project Genie, AMI Labs and AMD's purchase of World Labs

Google's Project Genie reached US Ultra subscribers on 29 January; game stocks fell the same day (Unity by more than 20%, Roblox by 12 to 13%, Take-Two by 10%). Runway launched its GWM-1 world model in December and raised $315 million at a $5.3 billion valuation in February. LeCun's AMI Labs raised $1.03 billion in March to build world models that do not generate pixels. On 28 September AMD agreed to buy World Labs, Fei-Fei Li's spatial-intelligence company, for about $8.2 billion in stock, with Li becoming AMD's chief scientist. Chip vendors are moving up the stack because world models are a large compute workload. Kling 4.0 (30-second clips, ten keyframes) entered early access the same day.

How video generation works

Video is hard because it multiplies the cost of an image by time, and every frame has to agree with every other.

Why video is harder than an image

A 1024-pixel-square image, compressed by a typical autoencoder, becomes about 4,100 tokens for a transformer. A five-second 720p clip at 24 frames per second, compressed by a Wan-style video autoencoder, becomes about 108,000 tokens, roughly 26 times more, and because attention cost grows with the square of the length, about 700 times more attention work per layer (the chapter's own arithmetic from published compression ratios). Consistency adds a second difficulty. Identity, lighting and physics have to hold across every frame, and viewers notice flicker.

Four building blocks

  • Video diffusion. The model learns by corrupting real clips with noise and learning to remove it. Generation starts from pure noise and cleans it up step by step, steered by a text embedding. Early systems were cascades of models; later ones do it in one pass.
  • Latent video autoencoders. Raw video is too big, so an autoencoder squeezes it into a small grid, diffusion runs there, and a decoder paints pixels back. Video versions compress time as well as space, and Wan's is 4× in time and 8×8 in space. "Causal" means a frame sees only earlier frames, so one latent space can hold both a still image and a video, and long clips can be streamed.
  • Spacetime patches and DiT. A transformer reads the latent video as a sequence of small 3D blocks, which lets one model train on clips of any length and shape.
  • Temporal attention. Frames exchange information either cheaply (spatial attention inside each frame plus temporal attention across frames) or fully (attention over every token, costlier but better at consistency). Google's Lumiere took a third route. It shrinks time and generates the whole clip at once, where keyframe-based systems can drift apart.

Sound, speed and interactivity

  • Joint audio and video. The older way made silent video and added sound afterward. The newer way denoises both together so lip movement and footsteps are born aligned. LTX-2 is the open blueprint, with a 14-billion-parameter video stream and a 5-billion-parameter audio stream linked by cross-attention. Veo 3 and Sora 2 internals are unpublished.
  • Flow matching and distillation. Flow matching teaches the network to move from noise to data along near-straight paths, so fewer steps are needed. Distillation trains a student to match a many-step teacher in one to four steps. CausVid cut 50 steps to 4, and Seedance 1.0 reports about a tenfold speedup.
  • Streaming and interactive worlds. Instead of denoising a whole clip, the model generates frame by frame, conditioned on past frames and a user's actions. The hard problem is drift. Training sees clean history, but deployment sees the model's own errors, and fixes include training on the model's own rollouts (Self Forcing).
  • Gaussian splats. The persistent-world format. A scene is millions of explicit 3D blobs with no network inside, so it can be saved, edited and exported. Marble exports splats and meshes.

Lesser-known stories about video

Eleven stories from the research notes, from the three founders of Runway to Fei-Fei Li's move to AMD.

  • A Chilean, a Greek and a Chilean founded Runway in a university art programme, and the company then helped create Stable Diffusion. Its ex-researchers founded Black Forest Labs.
  • Sora's artists leaked it. In November 2024 beta testers protested unpaid labour by posting access publicly; it was revoked in about three hours.
  • Air Head took heavy post-production. The film's 300-to-1 footage ratio and hand rotoscoping show what demo reels leave out.
  • The model that could not hold a pan. In Air Head, requested camera pans landed about 60% of the time. By 2026 camera movement is an explicit input in Flow, Runway's world model and Marble.
  • Nvidia's Cosmos used a lifetime of video each day. Its paper reports about 20 million raw hours, 100 million clips and 10,000 H100 GPUs for three months.
  • Meta launched a feed while Movie Gen stayed unreleased. Movie Gen, trained on up to 6,144 H100s with about 100 million videos, was never released as far as the research found; Vibes launched five days before Sora's app.
  • Project Genie hit game stocks. The launch knocked 10 to 24% off them in one day.
  • DeepMind argued both sides on physics. Its Physics-IQ benchmark (January 2025) found that video realism is unrelated to physical understanding; its Veo 3 paper (September 2025) argued video models are becoming zero-shot learners and reasoners. An ICML 2025 paper found that models memorize individual cases without learning the underlying laws.
  • Kling's architect left. Zhang Di reportedly moved to Alibaba in late 2025, and an anonymous model called HappyHorse-1.0 topped the Artificial Analysis ranking in April 2026, which Alibaba claimed (unverified).
  • Seedance in 72 hours. The studios reacted within 72 hours, and the speed of the reaction was itself news.
  • Fei-Fei Li's path from W.A.L.T to AMD. She co-authored W.A.L.T, founded World Labs with NeRF's Ben Mildenhall, and is now AMD's chief scientist after an $8.2 billion deal.

Are video models world simulators?

Sora's report, Google's Veo 3 paper and Genie 3 point toward video models learning how the world works, while LeCun and DeepMind's Physics-IQ benchmark point the other way, and nothing so far settles it.

The case for

Sora's report claimed that 3D consistency and object permanence emerged with scale, with admitted failures such as shattering glass. Jim Fan called it a data-driven physics engine. Google's Veo 3 paper found video models solving tasks they were never trained on. Genie 3 holds a world together for minutes and responds to user actions.

The case against

LeCun argues that predicting pixels wastes capacity on irrelevant detail and the right target is an abstract representation, as in V-JEPA 2. Physics-IQ found realism did not imply physical understanding. Runway's own team has admitted its Gen-4.5 still fails at causal reasoning and object permanence. Genie 3's memory lasts only minutes.

What is unresolved

Three things are unsettled. One is whether scale yields rules or only memorized cases, another is whether pixel space is the right space, and the third is whether persistence needs explicit 3D (Marble) or a learned state. OpenAI's decision to kill the Sora product while keeping its world-simulation research cuts both ways. The practical test for the next year is whether any of these systems can train a robot to do something it could not do before.

Money, law and open weights

Veo's API price per second fell from $0.75 to $0.40, Runway and Luma each raised more than $300 million, and a federal judge let Disney's suit against MiniMax go forward.

Prices and money

Veo 3's API launched in July 2025 at $0.75 per second of video with audio. By autumn 2026 Google lists Veo 3.1 at $0.40 per second standard and cheaper tiers at $0.05 to $0.12; Gemini Omni Flash is about $0.10 at 720p. Runway raised $308 million in April 2025 and $315 million at $5.3 billion in February 2026, with Nvidia and AMD among investors. Kuaishou's Kling reported annualized revenue of about $240 million at the end of 2025 and was reported at about $500 million by March 2026. It was also reported to be seeking funds for a planned 2027 Hong Kong listing. Luma raised $900 million in November 2025. MiniMax listed in Hong Kong in January 2026.

Open versus closed

The open models are Wan 2.1 and 2.2, HunyuanVideo, LTX-2 and Mochi. The closed ones are Seedance, Kling, Veo, Omni and Hailuo. Chinese labs shipped fastest and several released open weights, while US labs hold most of the capital and compute.

Law and provenance

A federal judge denied MiniMax's motion to dismiss Disney's suit in May 2026, studios sent cease-and-desist letters over Seedance, and Disney's OpenAI deal collapsed. Moonvalley's Marey, a licensed-data model, went public in July 2025, and Lionsgate signed a custom-model deal with Runway. On provenance, Google's SynthID had marked more than 10 billion items by May 2025, but watermarks are easy to strip and open models cannot be forced to mark outputs. xAI's Grok Imagine, launched in July 2025 with a "Spicy" mode, became a deepfake scandal that drew regulator action in several countries.