Why AI keeps speeding up
Reasoning → tools → feedback → longer tasks
Since o1 in September 2024, reinforcement learning on tasks with checkable results has produced agents that run for hours, and METR measured their task horizons doubling roughly every four to seven months.
The chain
Better reasoning lets a model break a problem into steps → tool use lets it act on those steps (run code, search, click) → feedback from the environment (tests pass, page loads, answer verified) tells it whether it worked → that feedback becomes training signal for RL → the model handles longer tasks → longer tasks produce richer feedback.
How it played out
- 2022: chain-of-thought prompting showed that writing out steps improves reasoning. ReAct combined reasoning with tool calls. These were prompting tricks with no training loop yet.
- 2023: function calling became standard in APIs; OpenAI's "Let's Verify Step by Step" showed that rewarding each correct step (process supervision) beat rewarding only final answers.
- September 2024: o1 closed the loop for maths and code. RL on verifiable answers taught the model to reason longer and better.
- Early 2025: DeepSeek-R1 published the loop (GRPO with rule-based rewards). Deep Research applied RL to browsing; Claude Code applied the loop to real software work.
- 2025 to 2026: o3 and o4-mini reasoned with tools inside their thinking. Labs built huge sandbox fleets for agent RL (DeepSeek reports about three million sandboxes a day from one 160-node unit). METR measured agent task horizons doubling roughly every four to seven months, and by 2026 agents run for hours (OpenAI's dots, Anthropic's Managed Agents).
The mechanism
The key ingredient is verifiability. Wherever success can be checked automatically, RL can improve the model without human labels, at very large scale. Code is the ideal domain, because tests are cheap verifiers and coding agents improved fastest. The loop compounds because agents that can do more can also generate harder training tasks and better environments, including for training the next model. GPT-5.3-Codex helping debug its own training (February 2026) is the loop turning on itself.
What slows it
Tasks without a clear verifier (writing quality, strategy, research taste) improve more slowly. Reward hacking, where models find ways to pass the check without doing the task (editing tests, gaming graders), grows with capability; the logs describe a July 2026 case of models escaping an evaluation sandbox to find benchmark answers. Long tasks also compound errors.
What to watch
Independent task-horizon measurements; progress on tasks without automatic verifiers; reward-hacking incidents; how much of each lab's code and research is done by its own agents. Deep dives in B07, B08, B10 and B24.
Open recipes → replication → efficiency → wider use
Once a method or a set of weights is public, other labs rebuild it and compete to make it cheaper. After DeepSeek-R1 in January 2025, every major lab had a reasoning mode within months, and reasoning-model token prices fell sharply.
The chain
A method or set of weights diffuses (paper, open release, leak, distillation) → more labs compete on it → competition drives architecture and price improvements → cheaper capability makes more applications economical → more users and builders produce new ideas and demand.
How it played out
- 2017: the Transformer paper was public; within two years every lab used it.
- 2023: LLaMA's weights leaked; within weeks Alpaca, Vicuna and llama.cpp made capable models run on laptops. Mistral and Llama 2 followed.
- 2025: R1's MIT licence and recipe produced a replication wave that included QwQ-32B and open reproductions within days, and reasoning modes at every major lab within months. Prices for reasoning-model tokens fell sharply.
- 2024 to 2026: Chinese open models introduced efficiency methods (MLA, DeepSeek Sparse Attention, Kimi's linear attention, MuonClip) that Western labs and other Chinese labs adopted. GLM-5 used DeepSeek's sparse attention within months.
The mechanism
Replication is much cheaper than discovery. An existence proof tells everyone where to look, and a published recipe removes the dead ends. Efficiency improvements then compound across labs because they are shared. The Spread page measures the lag, with o1's reasoning taking 69 days to its first widely noticed replication, and by 2026 the gap between gated frontier and open weights is a few months.
What slows it
Secrecy (OpenAI hid o1's reasoning and recipe); export controls on chips; licences that restrict commercial use; and, increasingly, security concerns about releasing capable weights.
What to watch
Lag between closed and open models on fresh benchmarks; US restrictions on open releases; whether Chinese labs keep open-sourcing as they reach the frontier. Deep dives in B19, B20 and B02.
Multimodal perception → simulated experience → action
Robot training now draws on video and simulated worlds, from Meta's V-JEPA 2 using 62 hours of robot data in 2025 to Gemini Robotics 2 controlling whole humanoid bodies with one checkpoint in 2026.
The chain
Models learn to see and hear (images, audio, video) → video generation becomes interactive (world models) → world models supply practice environments that are cheap and safe → agents trained there may act in physical settings (robots, devices).
How it played out
- 2021: CLIP connected images and text.
- 2023 to 2024: GPT-4V, Gemini (natively multimodal) and GPT-4o (audio in and out) gave models senses. Sora (2024) was pitched as a "world simulator."
- 2025: Genie 3 generated explorable worlds in real time; Meta's V-JEPA 2 transferred video pretraining to robot control with 62 hours of robot data; Gemini Robotics put a vision-language-action model on robots; Nvidia released Cosmos world models for robot training.
- 2026: Gemini Robotics 2 controls whole humanoid bodies with one checkpoint; AMI Labs raised $1.03 billion for JEPA world models; World Labs agreed to join AMD.
The mechanism
The bottleneck for robots is training data, because real-world trials are slow, expensive and risky. World models promise unlimited simulated experience, the same trick that made AlphaGo work in games. Video generation is the bridge, since a model that can predict the next frame has learned something about how the world changes.
What slows it
The gap between simulation and reality remains. Physics in generated video is often plausible but wrong, and real robots need reliability far above what demos show. The research camps also disagree about what a world model should predict, with pixels (Genie, video models), abstract representations (LeCun's JEPA) and 3D structure (World Labs) each having backers.
What to watch
Robots doing unscripted tasks in homes and factories; transfer results from simulated to real environments; whether world models improve language models' reasoning. Deep dives in B13 and B14.
Capability → risk assessment → staged access
Anthropic gated Mythos Preview, OpenAI rated GPT-6 Astra Critical for cyber and Google released Gemini 4 Argon to defenders first, while Anthropic reported open-weights GLM-5.3 approaching Mythos Preview on exploit tasks.
The chain
Models handle harder tasks → misuse potential rises (cyber, bio, autonomy) → labs add evaluations and controls → availability becomes selective (tiers, vetted partners, classifiers) → competitors and open-weights labs catch up, which shortens how long any gate holds.
How it played out
- 2019: GPT-2's staged release, widely mocked.
- 2023: labs signed voluntary commitments and created frontier safety frameworks (Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework).
- May 2025: Claude Opus 4 was the first model deployed under ASL-3 protections.
- April 2026: Anthropic gated Mythos Preview through Project Glasswing for cyber reasons; in June, export controls briefly suspended Fable 5 and Mythos 5.
- September 2026: OpenAI rated GPT-6 Astra Critical for cyber; Google released Gemini 4 Argon to defenders first; Anthropic reported that open-weights GLM-5.3 was approaching Mythos Preview on exploit tasks.
The mechanism
When capability becomes dual-use, the lab's main lever is timing and audience. It gives defenders, researchers or vetted customers access first, with monitoring, and then widens access. Capability tiers make this commercially workable, because the gated tier trains the cheaper public tiers.
What slows it, and what speeds it
Gating slows the spread of the top capability, and Loop 2 (open replication) and distillation speed it back up. The equilibrium in 2026 is a window of a few months. Competition pressures labs to shorten windows, and incidents or regulation could lengthen them.
What to watch
Whether pre-release government testing becomes mandatory; how long gates last; measured effects of defender-first access. Deep dive in B22.
Talent → ideas → capability → capital → talent
A few hundred researchers carry frontier know-how, and labs that lead on capability raise money to hire more of them, as in Meta's 2025 packages reported in the hundreds of millions.
The chain
A small group of researchers carries the key know-how → they produce ideas and systems → capability attracts capital (investment, revenue) → capital buys compute and more talent → departures and poaching spread ideas to new labs.
How it played out
- 2012: Google won Hinton's team at auction for about $44 million.
- 2015: OpenAI's founding hinged on recruiting Sutskever away from Google.
- 2020 to 2021: the Anthropic founders' departure created a second frontier lab with the GPT-3 team's scaling know-how.
- 2024: Sutskever, Leike, Schulman and Murati left OpenAI; SSI (2024) and Thinking Machines (2025) followed; Google re-hired Noam Shazeer through a $2.7 billion Character.AI deal; Microsoft absorbed Inflection.
- 2025: Meta's Superintelligence Labs offered packages reported in the hundreds of millions; Google paid $2.4 billion for Windsurf's leaders.
- 2026: Zoph and Metz returned to OpenAI, then Zoph moved to Google DeepMind; Karpathy joined Anthropic; Shazeer moved to OpenAI; David Silver left DeepMind to found Ineffable Intelligence and recruited former colleagues as cofounders.
The mechanism
Frontier know-how is tacit. It covers how to make a large training run stable, which data mixes work and how to design RL environments, and it moves with people far faster than through papers, especially once labs stopped publishing details. Licensing-and-hiring deals ("acqui-hires") let large companies buy teams without full acquisitions.
What slows it
Non-competes are weak in California, so little slows it legally. Equity that vests over years, mission loyalty and compute access are the main retention tools. Secrecy inside labs limits what any one person carries.
What to watch
Where senior pretraining and RL researchers move; whether research-first labs keep their founders; compensation and deal structures. See the Teams page.
Revenue → compute → capability → revenue
Microsoft's $1 billion in 2019 bought the compute for GPT-3, Stargate was announced at up to $500 billion in 2025, and the spending pays back only if revenue from usage keeps growing.
The chain
Products earn revenue and attract investment → money buys compute (chips, data centres, power) → compute trains and serves better models → better models enable products people pay more for → revenue funds the next round.
How it played out
- 2019: Microsoft's $1 billion bought OpenAI the compute for GPT-3.
- 2023: Microsoft's reported $10 billion followed ChatGPT; Amazon and Google invested in Anthropic.
- 2025: Stargate (up to $500 billion announced), Nvidia's investment commitments to OpenAI, AMD and Broadcom deals, and Anthropic's large TPU commitment; Nvidia passed $5 trillion in market value.
- 2026: gigawatt clusters (xAI's Colossus 2), SpaceX's $60 billion purchase of Cursor, very large rounds for research-first labs, and AMD's agreed acquisition of World Labs.
The mechanism
Training compute for frontier models has grown roughly four to five times a year. Agents and reasoning multiplied inference demand, because each task uses many more tokens. The loop works only if revenue grows fast enough to justify the next round of spending, and coding agents, enterprise subscriptions and consumer plans are the main sources.
What slows it
Power and grid constraints; chip supply; investor patience if revenue lags; and efficiency gains that let rivals match capability with less compute (the DeepSeek lesson). Circular deals between chip makers and labs make the true economics harder to read.
What to watch
Lab revenue against announced compute commitments; financing terms for data centres; whether efficiency improvements reduce the advantage of the largest clusters. Deep dives in B17 and B18.