The AI landscape

How the forces fit together

Capability brings adoption, adoption pays for compute, compute produces the next capability, and open release and safety rules speed up or slow down that cycle.

The forces work as a flywheel with brakes. Capability improvements (reasoning, tool use, longer context) make new applications possible. Adoption of those applications produces revenue and usage data, which justifies more compute. More compute, used more efficiently, produces the next round of capability. The open ecosystem speeds the whole loop up for everyone by spreading methods and weights, and puts price pressure on the leaders. Safety, rules and trust act as the brakes and steering. They decide which capabilities reach which users and on what terms.

Two features set 2025 to 2026 apart from earlier years. First, the flywheel now spins inside a few very large organizations (OpenAI with Microsoft and SoftBank, Google, Anthropic with Amazon and Google, Meta, xAI with SpaceX, and Alibaba, ByteDance and DeepSeek in China), so capital concentration matters as much as research. Second, the brakes have become part of the product. Staged releases, capability tiers and government pre-release testing now shape what "the best model" means for any given user.

The Connections page traces six loops step by step. The sections below take one force at a time.

Capability

Frontier models are now scored on finished tasks such as fixing software issues, running a terminal or a computer and producing professional work, where earlier tests scored answers to questions.

Where it stands

In October 2026 the frontier models (Anthropic's Fable and Mythos 5.1 tier and Opus 5.5, OpenAI's GPT-6 Astra and Sol, Google's Gemini 4 Argon and 3.x line, xAI's Grok 4.7, plus open-weights models such as Kimi K3, Qwen3.8, DeepSeek-V4 and GLM-5.x) are measured mainly on what they can do as agents. The tasks include resolving software issues end to end (SWE-bench Verified and Pro, DeepSWE), operating a terminal (Terminal-Bench) or a computer (OSWorld), browsing to find hard facts (BrowseComp) and completing professional work products (GDPval). Olympiad mathematics and competitive programming, open problems in 2023, have been at gold-medal level since mid-2025.

How we got here

Capability has come from three sources in turn. From 2018 to 2023 it came mainly from pretraining scale, meaning bigger Transformers on more text, following the scaling laws (GPT-2, GPT-3, GPT-4). From 2022 post-training made models usable, because instruction tuning and RLHF turned a text predictor into an assistant (InstructGPT, ChatGPT). From late 2024, reinforcement learning on verifiable tasks and test-time compute became the main lever. o1 and DeepSeek-R1 showed that training a model to think longer, rewarded by whether answers are correct, improved math, code and science sharply. In 2025 to 2026 that same RL machinery was pointed at agentic tasks, with real tools and environments, which produced the coding agents.

The mechanism

The current recipe has three stacked stages.

  1. Pretraining on tens of trillions of tokens of text, code, images and video builds broad knowledge and skills. It still matters, and Google argued with Gemini 3 that pretraining had headroom left.
  2. Mid-training and supervised fine-tuning shape the model toward reasoning traces, tool calls and long contexts.
  3. Large-scale reinforcement learning in environments with checkable outcomes (tests pass, answer correct, task completed) teaches the model to plan, try, check and correct. Labs now spend a meaningful fraction of their compute here, and DeepSeek reported RL budgets above 10% of pretraining for V3.2.

Test-time compute is the fourth lever. Letting a model think longer, or run several attempts in parallel and pick the best, raises accuracy at higher cost, and products expose this as effort settings.

Who leads

On most agentic benchmarks in the logs, the gated tiers of Anthropic, OpenAI and Google lead, with their public flagships close behind. Chinese open-weights models trail the closed frontier by a few months. xAI competes near the top on some evaluations. Meta returned to frontier-class models in 2026 with Muse Spark after the Llama 4 stumble.

The counter-argument

Benchmarks saturate and are partly gamed, and contamination and the choice of test scaffold move scores by many points. Agentic benchmarks are closer to real work but still curated. The strongest skeptical case, argued by researchers including Sutskever and LeCun, is that today's models generalize far worse than humans. They need enormous data and still fail on simple novel problems (ARC-style puzzles until recently, and the newer ARC-AGI-3, where GPT-6 Astra scored 62.7% in September against humans' 100%). On that view, capability is broad but brittle, and the next leap needs new ideas.

What to watch

Independent long-horizon reliability measurements (METR-style task horizons), results on fresh benchmarks released after a model's training, and whether continual-learning methods let models improve from deployment experience.

Adoption

Coding agents lead paid adoption, with Cursor and Cognition each past $1 billion in annualized revenue, and regulated fields such as medicine and law are moving more slowly.

Where it stands

Consumer chat assistants are mass-market products. ChatGPT reported hundreds of millions of weekly users through 2025, and Gemini, Meta AI, Claude, Doubao and DeepSeek have very large user bases. The fastest business adoption is in software. Coding agents (Claude Code, Codex, Cursor, Copilot, Devin) became the clearest proof of value, with Cursor passing $1 billion in annualized revenue in late 2025 and Cognition reporting the same in September 2026. In 2026 the focus moved to general knowledge work through agent products (Claude Cowork, OpenAI's ChatGPT agent and dots, Meta's Muse agent).

How we got here

Adoption reached each field in the order its results could be checked. Code came first because results can be checked automatically, which helps both the model's training and the customer's trust. Customer support, document drafting and research synthesis followed. Regulated and physical domains (medicine, law, manufacturing) move more slowly, through specialized products and certifications.

The mechanism

Adoption depends on integration, which means access to the organization's data and tools (connectors, MCP), permissions and audit trails, predictable cost, and someone accountable for checking output. That is why 2025 to 2026 product launches are mostly agent frameworks, marketplaces and managed agents, with few new chat features, and why labs now sell implementation help and training programmes alongside models.

Who leads

OpenAI leads in consumer reach; Anthropic built its business on developers and enterprises, especially coding; Google leads in distribution through Search, Android and Workspace; Microsoft through Office and GitHub. In China, ByteDance's Doubao and Alibaba's Qwen apps compete for consumers while DeepSeek's low prices shaped developer adoption.

The counter-argument

High usage does not prove higher productivity. A 2025 METR randomized trial found experienced open-source developers were slower with AI tools on their own repositories, despite believing they were faster. Many enterprise pilots fail to reach production. Revenue growth is real but concentrated in a few products, and much of it is other AI companies paying each other.

What to watch

Measured outcomes from large deployments; pricing moves from per-token to per-task or per-seat; labour-market data on occupations most exposed to agents.

Compute and cost

Frontier training compute has grown about four to five times a year, while the cost of a given level of capability has fallen even faster.

Where it stands

The frontier labs and their backers are building data centres at gigawatt scale. The projects include OpenAI's Stargate programme with SoftBank and Oracle (announced January 2025 at up to $500 billion), xAI's Colossus 2 (reported as the first gigawatt training cluster in January 2026), Meta's Prometheus and Hyperion, Microsoft's Fairwater sites and Anthropic's large commitments to Google TPUs and Amazon Trainium. Nvidia remains the dominant supplier, Google's TPUs are the strongest alternative, and AMD gained ground with OpenAI and, per the logs, by buying World Labs. Power and grid connections are now the binding constraint in many places.

How we got here

AlexNet ran on two gaming GPUs in 2012. GPT-3 trained on a 10,000-GPU Microsoft supercomputer in 2020. GPT-4-class models used tens of thousands of A100s. Each generation since has increased training compute by roughly four to five times a year (Epoch AI's estimate of the long-run trend). Meanwhile the cost of a given capability fell even faster. GPT-4-level output that cost $30 to $60 per million tokens in 2023 cost cents by 2025.

The mechanism

Total spend and unit cost move in opposite directions. Total spend rises because each frontier run and each wave of agent usage needs more compute, and because reasoning and agents consume far more tokens per task than chat. Unit cost falls because of better chips, lower precision (FP8, FP4), mixture-of-experts, sparse and linear attention, distillation into smaller models, and better serving software. DeepSeek's 2025 shock was a demonstration of the second curve under export-control pressure.

Who leads

US hyperscalers and labs have far more compute than anyone else. Chinese labs are constrained by export controls on Nvidia's best chips and have responded with efficiency engineering and a shift toward Huawei Ascend, visible in DeepSeek's September 2026 Ascend libraries and Meituan's model reportedly trained entirely on Chinese accelerators.

The counter-argument

The bubble case says capital spending far exceeds AI revenue today, some deals are circular (chip suppliers investing in customers who buy their chips), and the debt financing data centres could be fragile if growth slows. The efficiency case cuts both ways, because if capability gets much cheaper the largest clusters may earn less than planned.

What to watch

Lab revenue against compute commitments; power-plant and grid deals; next-generation chips (Nvidia's Rubin, Google's TPUs, AMD's MI series); Chinese domestic chip output.

The open ecosystem

Open-weights models, mostly from Chinese labs, trail the closed frontier by a few months, and replicating a new capability now takes weeks where it once took years.

Where it stands

By 2026 the strongest open-weights models are mostly Chinese. They include Moonshot's Kimi K3 (2.8 trillion parameters, July), Alibaba's Qwen3.8 (2.4 trillion, open in August), DeepSeek-V4, Zhipu's GLM-5 line and Xiaomi's MiMo-V2.6 (top-rated open model at launch in September). American open models include OpenAI's gpt-oss (2025), Google's Gemma 4 and, four months after going closed, Meta's Muse Glimmer 30B (August 2026). Open models typically trail the closed frontier by a few months.

How we got here

GPT-2's staged release in 2019 framed openness as a safety question. Meta's LLaMA (February 2023), and its leak, created the modern open ecosystem, and Llama 2 and Mistral made it commercial. DeepSeek-R1 (January 2025) was the turning point. It was a frontier-class reasoning model with an MIT licence and a published recipe, and it pulled reasoning into every lab's roadmap within weeks. Through 2025 Chinese labs released faster and more openly than US labs, and their downloads overtook American ones.

The mechanism

Openness speeds diffusion in three ways. Published papers spread ideas (GRPO, MLA, sparse attention), open weights let anyone fine-tune, distil and deploy, and open tooling (vLLM, llama.cpp, Hugging Face) makes running models cheap. The Spread page shows the result, with the lag between a frontier capability and its first replication falling from years to weeks. Distillation, sometimes contested, is a fourth channel.

Who leads

Alibaba's Qwen family has the largest ecosystem of derivative models; DeepSeek and Moonshot set the technical pace; Zhipu and MiniMax compete on agentic coding. Hugging Face is the distribution hub.

The counter-argument

Open weights cannot be recalled. Anthropic's September 2026 study of GLM-5.3's exploit ability, if it holds up, shows offensive cyber capability now available in a downloadable model. The national-security case against open frontier weights has grown stronger as capability rises. There is also a business question, since some open releases are strategic and aim at commoditizing a competitor's product.

What to watch

Whether US policy restricts open-weights releases above a capability threshold; whether Chinese labs keep releasing their best models openly as they approach the frontier; whether Meta's return to open models continues.

Science, media and physical AI

Models now work in biology and mathematics, generate video clips of up to 30 seconds with sound, build explorable worlds and drive humanoid robots.

Where it stands

  • Science includes AlphaFold's protein-structure work, which won a 2024 Nobel Prize. AI systems discover algorithms (AlphaEvolve, 2025), reach gold level in olympiad maths, and in 2026 run large agent swarms over biological databases (Anthropic's enzyme discovery in September). Microsoft, Google and Anthropic all launched life-sciences programmes.
  • Media includes video models that generate clips of up to 30 seconds with synchronized sound (Veo, Seedance 2.x, Kling 4.0, Wan 3.0). Voice models are close to indistinguishable from humans (Eleven v4), and music generators moved from lawsuits to licensing deals. OpenAI shut down Sora in 2026, and media generation remains expensive and legally contested.
  • World models and robots include Genie 3, which generates explorable worlds in real time, and World Labs, which builds persistent 3D worlds and is joining AMD. Gemini Robotics 2 drives humanoid bodies, and LeCun's AMI Labs bets on JEPA-style world models.

How we got here

Science applications started with DeepMind's games-to-science strategy (AlphaGo, then AlphaFold). Media generation went from GANs to diffusion (2022) to native audio-video (2025). World models grew out of game environments and video generation, and attracted heavy funding in 2025 to 2026 as an alternative or complement to language models.

The mechanism

Progress is fastest in these domains where outcomes can be simulated or checked (protein structure, proofs, code, physics in a simulator) and slowest where the real world must be tested (wet labs, clinical trials, robots in homes). AI-for-science increasingly means agents that propose and rank hypotheses before expensive experiments.

Who leads

Google DeepMind leads in science breadth and world models; ByteDance, Kuaishou, Alibaba and Google lead video; ElevenLabs leads voice; robotics is split between Google, Nvidia, Physical Intelligence, Figure, Tesla and Chinese humanoid makers.

The counter-argument

Many "AI discovered X" claims are later qualified, because the hard parts (validation, manufacturing, regulation) remain human and slow. Video models produce plausible footage without reliable physics. Robotics demos are curated. Claims of world understanding should be tested against transfer to new environments.

What to watch

Peer-reviewed discoveries with experimental confirmation; robotics deployments outside labs; licensing settlements that define what media models may train on.

Rules, safety and trust

The top tier at every major lab now ships in stages, to vetted cyber defenders or partners first, and governments run voluntary pre-release testing.

Where it stands

The top tier of every major lab is released in stages, with vetted cyber defenders or partners first (Anthropic's Mythos via Project Glasswing, OpenAI's GPT-6 Astra rated Critical, Google's Gemini 4 Argon). Labs publish system cards and safety frameworks (Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, Google DeepMind's Frontier Safety Framework). Governments run voluntary pre-release testing; the EU AI Act's obligations for general-purpose models are phasing in; California's SB 53 requires transparency from frontier developers; and the US has used export controls on chips and, per the logs, briefly on a model.

How we got here

Safety began as a research agenda (Concrete Problems in AI Safety, 2016; RLHF as alignment, 2017), became a founding rationale (Anthropic, 2021), turned public in 2023 (the pause letter, Hinton leaving Google, the Bletchley summit), and became operational in 2024 to 2025 as labs adopted capability thresholds and the first models were treated as crossing them (Claude Opus 4 under ASL-3 in May 2025). Research on deceptive behaviour (sleeper agents, alignment faking, scheming evaluations, chain-of-thought monitorability) moved from hypothetical to measured.

The mechanism

Labs work from capability thresholds plus safeguards. They test each new model for dangerous capabilities (cyber, bio, autonomy), and if it crosses a threshold they add classifiers, access controls, monitoring or delayed release. Customer trust is built in a similar way, with audit logs, data controls and sandboxed agents.

Who leads

Anthropic built its identity around safety and pioneered tiered release; OpenAI and Google DeepMind have converged on similar practices. Independent evaluators (METR, Apollo Research, the UK AI Security Institute, the US CAISI) test models before release.

The counter-argument

Critics on one side say voluntary frameworks are self-graded and bend under competitive pressure, and this year's public resignations from Anthropic and OpenAI make that argument. Critics on the other side say gating protects incumbents and slows useful technology. Fast open replication (Force 4) limits how long any gate can hold.

What to watch

Whether pre-release testing becomes mandatory in the US; how the EU enforces its general-purpose AI rules; whether chain-of-thought stays readable as labs optimize harder; and whether a serious AI-enabled incident changes the politics.