How we got here

2012 to 2014

AlexNet wins ImageNet, and Google buys DNNresearch and DeepMind

Hinton's Toronto team showed in 2012 that a big neural network trained on GPUs could beat decades of hand-built computer vision, and the biggest tech companies then bought or hired the people who knew how.

AlexNet wins ImageNet in 2012

For most of the 2000s, neural networks were an unfashionable corner of AI. The mainstream used hand-engineered features and statistical models, while a small group around Geoffrey Hinton in Toronto, Yann LeCun at NYU and Yoshua Bengio in Montreal kept arguing that networks with many layers would win if they had enough data and compute.

In September 2012 they got their proof. Alex Krizhevsky, Ilya Sutskever and Hinton entered a convolutional network, later called AlexNet, in the ImageNet image-recognition challenge. It cut the top-5 error rate to about 15%, against about 26% for the next-best entry. The model was trained on two consumer Nvidia gaming cards (GTX 580s). Two ingredients that would define the next decade were already present. One was a large labelled dataset (ImageNet, assembled by Fei-Fei Li's group) and the other was GPUs repurposed from graphics to matrix math.

Google wins the DNNresearch auction

The result set off a buying spree for researchers. In December 2012 Hinton and his two students set up a tiny company, DNNresearch, and auctioned themselves to tech companies from a hotel room during the NIPS conference at Lake Tahoe. Baidu, Microsoft, Google and the then-small DeepMind bid, and Google won for about $44 million. Sutskever moved to Google Brain, the deep-learning group Jeff Dean and Andrew Ng had started in 2011.

That auction set a pattern that runs through this whole story. Frontier AI capability lives in a few hundred researchers, and labs compete for them as aggressively as for chips. Facebook hired LeCun to found FAIR in late 2013. Google bought DeepMind in early 2014.

DeepMind's Atari result and the sale to Google

DeepMind had been founded in London in 2010 by Demis Hassabis, Shane Legg and Mustafa Suleyman, with a mission stated in the language of artificial general intelligence when that phrase was still embarrassing in academia. Its first famous result, in late 2013, was a single network (DQN) that learned to play seven Atari games from pixels and the score alone, using reinforcement learning (RL), which learns from trial, error and reward and needs no labelled examples.

Google bought DeepMind in January 2014 for a reported $400 to $650 million, after Facebook also tried. Hassabis secured unusual terms, an ethics and safety board and a degree of independence that would be fought over for years. Mallaby's The Infinity Machine and Olson's Supremacy start the Hassabis thread here. He was a games prodigy and neuroscientist who believed the route to general intelligence ran through games, then science.

Word vectors, translation and attention

Four quieter papers from this period matter more for today's AI than the headline results.

  • word2vec (2013) showed that words could be represented as vectors whose geometry captures meaning.
  • Sequence-to-sequence learning (Sutskever, Vinyals and Le, 2014) used recurrent networks to map one sentence to another, the basis of neural translation.
  • Attention (Bahdanau, Cho and Bengio, 2014) let a translation model look back at the relevant input words instead of squeezing a whole sentence into one vector. Attention is the seed of the Transformer.
  • GANs (Goodfellow, 2014) started the generative-image line that later handed over to diffusion.

The constraint was still that recurrent networks process text one word at a time, which made them slow to train on very large datasets, and removing it is the next era's breakthrough. See B01 and B12 for deep dives.

2015 to 2016

OpenAI is founded and AlphaGo beats Lee Sedol

OpenAI was founded in December 2015 as a counterweight to Google and DeepMind, and in March 2016 DeepMind's AlphaGo beat Lee Sedol 4 to 1 and showed the public that learning systems could find ideas humans had missed.

Why OpenAI was founded

By 2015 some technologists worried about concentration, because Google appeared to own most of the field's top people. Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever and a small founding group announced OpenAI in December 2015 as a non-profit with $1 billion in pledges (far less was actually contributed in the early years). The stated mission was to ensure that artificial general intelligence benefits all of humanity, with research shared openly.

Recruiting Sutskever away from Google was the decisive hire. Musk later described it as one of the hardest recruiting battles he had fought. Karpathy, Schulman, Zaremba and others joined the founding team. Hao's Empire of AI and Hagey's The Optimist both treat this founding as the origin of a tension that never went away. The mission was framed around safety and openness, and it was funded and led by people who also wanted to win a race.

AlphaGo and Move 37

In March 2016 DeepMind's AlphaGo beat Lee Sedol, one of the world's strongest Go players, 4 to 1 in Seoul. Go had been considered a decade away because the number of possible positions makes brute-force search hopeless. AlphaGo combined deep networks that judged positions and suggested moves with Monte Carlo tree search, and improved by playing against itself.

Two moves from the match became famous. Move 37 in game two was a play professionals initially read as a mistake and later called creative. Lee's Move 78 in game four, the only game he won, found a weakness the system had not seen. The technical lesson for the field was that learned intuition plus search plus self-play could exceed human expertise in a closed domain. The same combination of "think longer at decision time" and self-improvement returns in 2024 as reasoning models. See B15 for a deep dive.

Gym, robotics and Dota 2 at OpenAI

OpenAI's early work was mostly outside language. It released Gym (an RL toolkit) and Universe, worked on robotics (a robot hand that later solved a Rubik's cube) and built OpenAI Five, which beat professional Dota 2 players in 2019. These projects taught the lab how to run very large training jobs. It still lacked a single bet that could justify the money it would need.

2017 to 2018

Google publishes the Transformer, and OpenAI and Google train GPT and BERT

Eight Google researchers removed recurrence from sequence models in June 2017, producing the Transformer that every frontier model still builds on, and OpenAI bet its future on pretraining Transformers on raw text.

The Transformer

In June 2017 eight Google researchers (Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin) published "Attention Is All You Need." The Transformer dropped recurrence entirely, so every word attends to every other word in parallel. That made training massively parallel on GPUs and TPUs, which meant models could finally soak up far more data.

The paper presented a better translation model, and the architecture turned out to scale. Google used it widely, including in BERT, but did not turn it into a general assistant. By March 2024 all eight authors had left Google, most to found companies (Cohere, Character.AI, Adept, Essential AI, Sakana AI, Inceptive, NEAR). The "Google invented it, others shipped it" line starts here, though Google's internal language models (Meena, LaMDA) were capable and held back by caution about reputation and safety. See B01 for a deep dive.

GPT and BERT

Two families of Transformer language models appeared within months of each other in 2018.

  • GPT (OpenAI, June 2018) was trained to predict the next word on a large corpus of books, then fine-tuned for specific tasks. Alec Radford led the work.
  • BERT (Google, October 2018) was trained to fill in masked words using context from both sides. It was better on many benchmarks and went into Google Search a year later.

BERT won the benchmarks of its day. Next-word prediction won the decade, because a model that generates text can be asked to do anything in text, and because its training signal scales without limit.

The first RLHF paper, June 2017

A June 2017 paper by Paul Christiano, Jan Leike and colleagues at OpenAI and DeepMind showed that an agent could learn a task from a few thousand human judgments between pairs of short video clips. A reward model learns what people prefer, and RL optimizes against it. This is the origin of RLHF (reinforcement learning from human feedback), the method that later turned GPT-3 into ChatGPT. It was framed as alignment research, a way to give systems goals that are hard to write down. See B05 for a deep dive.

AlphaZero and Musk's exit

In December 2017 DeepMind's AlphaZero learned chess, shogi and Go from the rules alone through self-play, beating the best specialist programs within hours of training. At OpenAI, Musk left the board in February 2018 after disagreements over control and direction. The lab now needed a new source of very large amounts of money.

2019 to 2020

GPT-2, Microsoft's $1 billion, scaling laws and GPT-3

OpenAI created a capped-profit company, took Microsoft's $1 billion and compute, and showed with GPT-3 that scaling a Transformer to 175 billion parameters produced new abilities.

GPT-2 and "too dangerous to release"

In February 2019 OpenAI announced GPT-2, a 1.5-billion-parameter model that wrote surprisingly coherent paragraphs, and initially withheld the full model citing misuse risk. The staged release (full weights in November 2019) was mocked by some researchers as a publicity stunt and praised by others as a responsible-disclosure experiment. Either way, it put OpenAI's language models in the news and established the habit of treating release itself as a safety decision.

The capped-profit company and Microsoft

In March 2019 OpenAI created a "capped-profit" company controlled by the non-profit, so investors could earn up to a fixed multiple. Sam Altman became CEO. In July 2019 Microsoft invested $1 billion, largely in Azure compute, and built OpenAI a dedicated supercomputer. This structure, a non-profit board controlling a company raising commercial capital, is the root of the November 2023 crisis and of the long restructuring that ended in 2025.

Bill Gates is often remembered as an OpenAI funder, but the money came from Microsoft. Gates's role, covered in The Optimist and in Gates's own writing, came in 2022, when he became the skeptic OpenAI had to convince (see the 2022 era).

Scaling laws

In January 2020 Jared Kaplan, Sam McCandlish, Dario Amodei and colleagues at OpenAI published "Scaling Laws for Neural Language Models." Test loss fell as a smooth power law with model size, data and compute, over many orders of magnitude. The practical meaning was that you could predict, before spending the money, roughly how much better a ten-times-bigger model would be. That turned AI progress from a research gamble into something closer to a capital-allocation problem. Rich Sutton's 2019 essay "The Bitter Lesson" made the same argument in philosophical terms. General methods that use more compute beat clever hand-built knowledge. See B03 for a deep dive.

GPT-3

GPT-3, described in May 2020 and offered through an API from June, had 175 billion parameters trained on about 300 billion tokens. Its surprise was in-context learning. Shown a few examples in the prompt, it would do a new task without retraining. Microsoft took an exclusive license to the underlying model in September 2020. Inside OpenAI, GPT-3 confirmed the scaling bet, and it also sharpened a disagreement about how fast to commercialize and how much safety work should gate releases.

AlphaFold 2

In November 2020 DeepMind's AlphaFold 2 effectively solved the protein-structure prediction challenge at CASP14, predicting 3D shapes with accuracy competitive with experiments. It was the strongest evidence yet for Hassabis's thesis that AI's biggest payoff would be in science, and it won Hassabis and John Jumper a share of the 2024 Nobel Prize in Chemistry. See B16 for a deep dive.

2020 to 2021

Anthropic's founders leave OpenAI, and Codex and Copilot launch

Dario Amodei and a group of OpenAI's safety and scaling leaders left to found Anthropic, while OpenAI and Microsoft turned GPT-3 into Codex and GitHub Copilot, the first AI product developers used every day.

Why the Anthropic founders left

At the end of 2020 Dario Amodei, then OpenAI's vice president of research, left with his sister Daniela Amodei, Tom Brown (lead author on GPT-3), Jared Kaplan, Sam McCandlish, Jack Clark, Chris Olah and others. Anthropic announced itself in May 2021 with $124 million. They left after GPT-3 and before GPT-4.

The public account is a disagreement about priorities. The departing group believed scaling would keep working and would produce very powerful systems soon. Precisely because they believed that, they wanted a company organized around safety research, interpretability and careful deployment, with governance designed for it (Anthropic became a public benefit corporation and later added a Long-Term Benefit Trust). Reporting in Empire of AI and Supremacy adds friction over OpenAI's commercialization and the Microsoft relationship. Every book notes the irony that the split created a second frontier lab and intensified the race it was meant to slow.

Codex, Copilot and CLIP

OpenAI fine-tuned GPT-3 on public code to make Codex, which powered GitHub Copilot (technical preview June 2021). Copilot was the first widely used product built on a large language model, and it established coding as the domain where AI's value is easiest to measure, because code compiles and tests pass or fail. That property, verifiable output, returns in 2024 and 2025 as the engine of reasoning and agents.

In January 2021 OpenAI also released CLIP, which learned to connect images and text from 400 million captioned web images, and the first DALL·E. CLIP became the hidden component of the generative-image boom. See B09 and B11 for deep dives.

Why Google held back LaMDA

Through 2021 and early 2022 Google had among the most capable language models (LaMDA, then PaLM) and the most compute, via its TPU chips. It showed LaMDA conversing at Google I/O in May 2021 but did not release it broadly. The reasons were rational for an incumbent, since a chatbot that said something false or offensive could damage the search business and the brand. That caution is the background to Google's "code red" a year and a half later.

2022

InstructGPT, Stable Diffusion and the launch of ChatGPT

InstructGPT showed that RLHF made models useful to ordinary people, and DALL·E 2, Midjourney and Stable Diffusion brought image generation to the public. ChatGPT, launched as a research preview on 30 November, became what was then the fastest-growing consumer product on record.

InstructGPT teaches models to follow instructions

In January 2022 OpenAI published InstructGPT. Humans wrote example answers and ranked model outputs, a reward model learned those preferences, and RL tuned GPT-3 toward them. A 1.3-billion-parameter InstructGPT was preferred by labellers over the 175-billion-parameter GPT-3. The lesson was that how a model is post-trained can matter more to users than raw size. Anthropic published its parallel "helpful and harmless" RLHF work in April 2022 and Constitutional AI, where the model critiques itself against written principles, in December. See B05 for a deep dive.

Compute-optimal training and chain of thought

Two research results in early 2022 reshaped how models were built and used.

  • Chinchilla (DeepMind, March 2022) showed that most large models were undertrained, because for a fixed compute budget a smaller model trained on more data does better. The field re-balanced toward far more training tokens.
  • Chain-of-thought prompting (Google, January 2022) showed that asking a model to write out intermediate reasoning steps sharply improved math and logic. It was a prompting trick then and becomes a training target in 2024. See B07 for a deep dive.

DALL·E 2, Midjourney and Stable Diffusion

DALL·E 2 (April 2022), Midjourney's open beta (July) and Stable Diffusion (August) put high-quality image generation in public hands within five months. Stable Diffusion's open weights were the first big open-model release, and lawsuits from artists and Getty Images followed within months. Diffusion models, which learn to turn noise into images step by step, had overtaken GANs. See B11 for a deep dive.

Gates and the AP Biology test

In mid-2022 Bill Gates told the OpenAI team he would be impressed only when a model could pass the AP Biology exam, which takes reasoning as well as recall, and said he expected that to take years. In September 2022, according to Gates's own account, OpenAI showed him a model (the system that became GPT-4) answering 59 of 60 AP Bio multiple-choice questions and writing thoughtful answers to the open questions. Gates later wrote that it was the most important technology demonstration he had seen since the graphical user interface. The episode explains Microsoft's next move, since GPT-4 had already been trained by summer 2022 and Microsoft's leadership had seen it.

ChatGPT

On 30 November 2022 OpenAI released ChatGPT, a GPT-3.5 model tuned with RLHF for conversation, as a "low-key research preview." Accounts in Empire of AI and contemporaneous reporting describe a fast launch, partly to get ahead of rival chatbots. It reached about a million users in five days. The underlying science had changed little. A free chat box made the capability legible to everyone. Google declared an internal "code red."

2023

GPT-4, LLaMA's leak, the Google merger and Altman's firing

GPT-4 set a capability bar that took rivals a year to meet, open weights escaped from Meta, Google merged its two labs, and OpenAI's board briefly fired its CEO.

Microsoft's $10 billion and Bing's Sydney

In January 2023 Microsoft announced a multiyear investment reported at $10 billion. In February it launched a GPT-4-powered Bing chat that produced the "Sydney" persona in long conversations, declaring love, arguing and threatening. It was the first public lesson that deployed frontier models could behave in unexpected ways, and it showed how thin post-training safety layers could be.

GPT-4

GPT-4 (14 March 2023) passed professional exams at high percentiles, accepted images, and was far more reliable than GPT-3.5. OpenAI's technical report disclosed almost nothing about architecture, data or compute, citing competition and safety, a sharp break from the lab's open origins. It also reported that OpenAI had predicted GPT-4's final loss from much smaller runs, which supported the scaling story. Anthropic released Claude the same day, and Google rushed out Bard a week later. Gemini, Google's real answer, arrived only in December.

LLaMA leaks and open models spread

In February 2023 Meta released LLaMA to researchers, and within about a week the weights leaked publicly. Stanford's Alpaca showed in March that a cheap fine-tune on 52,000 instruction examples could approximate ChatGPT-style behavior. Meta then leaned in. Llama 2 (July 2023) came with a commercial licence, and Mistral, a Paris startup founded by ex-DeepMind and Meta researchers, released strong small models from September, then Mixtral, an efficient mixture-of-experts model, in December. From here on, every frontier capability gets an open-weights counterpart within months. See B19 and B02 for deep dives.

The pause letter, Hinton's exit and the Google Brain merger

The safety debate became public. A March "pause" letter asked labs to stop training beyond GPT-4 for six months (none did). Hinton left Google in May to speak freely about risk. A one-sentence statement equating AI extinction risk with pandemics and nuclear war was signed by lab leaders including Altman, Amodei and Hassabis. OpenAI launched a Superalignment team in July, promising 20% of its compute.

In April 2023 Google merged Google Brain and DeepMind into Google DeepMind under Hassabis, ending a decade of internal rivalry over compute and credit. Mallaby's book treats the merger as Hassabis finally getting the resources for the race he had warned about.

OpenAI's board fires Sam Altman

On 17 November 2023 OpenAI's non-profit board fired Sam Altman, saying he had not been consistently candid. Within five days nearly all employees threatened to leave for Microsoft, and Altman was reinstated with a new board. Sutskever, who had voted to remove him, reversed himself publicly. Reuters and The Information reported in the same week a project called Q*, said to improve mathematical reasoning, which was the first public footprint of what became o1. The crisis matters to the research story for two reasons. It showed that the non-profit structure could not actually constrain the company, and it set off a slow departure of safety-focused researchers. See B07 for a deep dive.

2024

Gemini 1.5, GPT-4o, the OpenAI departures and o1

Gemini 1.5 reached a million tokens of context, GPT-4o handled text, images and audio natively, OpenAI's safety team dispersed, and o1 showed that spending more compute on thinking could substitute for building a bigger model.

Long context and native multimodality

Gemini 1.5 (February 2024) handled a million tokens of context, enough for hours of video or an entire codebase, and showed near-perfect recall on "needle in a haystack" tests. OpenAI previewed Sora, a video model, the same day. GPT-4o (May) handled text, images and audio natively in one model, which made fluid voice conversation possible. Claude 3 (March) and Claude 3.5 Sonnet (June) moved Anthropic to the top of many coding evaluations, which became its commercial wedge. Meta's Llama 3.1 405B (July) was the largest openly released model to date. See B12 and B13 for deep dives.

Departures from OpenAI and the founding of SSI

Ilya Sutskever left OpenAI in May 2024, and Jan Leike, co-lead of Superalignment, resigned days later, saying safety culture had "taken a backseat to shiny products," and joined Anthropic. John Schulman followed to Anthropic in August, and Mira Murati left in September. In June Sutskever founded Safe Superintelligence Inc. (SSI) with Daniel Gross and Daniel Levy, with a single goal and, deliberately, no product along the way. SSI raised $1 billion at a $5 billion valuation in September 2024. Its thesis is that the next leap requires research as well as scale.

The scaling-wall debate

In late 2024 The Information and Bloomberg reported that the next generation of large pretraining runs at OpenAI, Google and Anthropic had shown smaller gains than hoped. Sutskever said at NeurIPS in December that "pretraining as we know it will end" because public text data is finite. Some framed this as "scaling is over." The next breakthrough, o1, showed that scaling had found a second axis, spending compute at inference time. See B04 for a deep dive.

o1 and reasoning models

On 12 September 2024 OpenAI released o1-preview. The model was trained with large-scale reinforcement learning to produce a long hidden chain of thought before answering, and its accuracy rose both with more RL training and with more thinking time at inference. On a 2024 math olympiad qualifier (AIME) the full o1 model went from about 12% for GPT-4o to 74% with one try (o1-preview scored about 45%). The ideas had a lineage in chain-of-thought prompting, process supervision ("Let's Verify Step by Step," 2023), and Noam Brown's work on poker and Diplomacy agents, where letting a system think longer at decision time was worth enormous amounts of training. OpenAI hid the reasoning and the recipe. In December it previewed o3, which scored 75.7% on the ARC-AGI-1 benchmark at a high cost per task. See B08 for a deep dive.

Nobel prizes, computer use and MCP

In October 2024 Hinton shared the Nobel Prize in Physics (with John Hopfield) for foundational work on neural networks, and Hassabis and Jumper shared the Chemistry prize (with David Baker) for protein structure. Later that month Anthropic released "computer use," letting Claude operate a screen with mouse and keyboard, and in November the Model Context Protocol (MCP), an open standard for connecting models to tools and data. Both were early infrastructure for agents. On 26 December DeepSeek released V3, an efficient mixture-of-experts model whose reported final training run cost about $5.6 million, and almost nobody outside the field noticed. See B09 for a deep dive.

Early 2025

DeepSeek releases R1 and other labs ship thinking models

DeepSeek, a Chinese lab, published R1 on 20 January 2025, an open reasoning model and recipe that matched o1 at a fraction of the price, and within weeks every frontier lab shipped its own thinking model.

R1-Zero and GRPO

DeepSeek, funded by the Hangzhou quant fund High-Flyer and led by Liang Wenfeng, released R1 on 20 January 2025 under an MIT licence. The paper's striking result was R1-Zero. Applying RL with simple rule-based rewards (is the math answer right, does the code pass) directly to the V3 base model produced long, self-correcting reasoning without any human-written reasoning examples. The model learned to pause, re-check and backtrack on its own, and the paper's authors called one such instance an "aha moment." The RL algorithm, GRPO, dropped the separate critic network PPO requires, which made RL cheaper. DeepSeek also distilled R1's reasoning into small open models anyone could run. See B08 and B20 for deep dives.

Nvidia's $589 billion drop and the $5.6 million figure

On 27 January the DeepSeek app topped the US App Store and Nvidia lost about 17% of its value in a day, roughly $589 billion, on fears that frontier AI would need far less hardware. The headline "$5.6 million" covered only the final pretraining run of V3 and excluded research, failed runs, staff and the GPU fleet behind it. Analysts estimated DeepSeek's hardware in the tens of thousands of GPUs. The durable lesson was efficiency under constraint. Export controls had limited DeepSeek's chips, so it engineered around memory and bandwidth (mixture-of-experts, multi-head latent attention, FP8 training). OpenAI and Microsoft said they were investigating whether DeepSeek had distilled from OpenAI's outputs.

The day after R1, President Trump announced Stargate, a $500 billion data-center venture led by OpenAI, SoftBank and Oracle. The two announcements pulled 2025 in two directions, toward lower cost and efficiency and toward raw capital.

How fast other labs copied o1

Once a recipe was public, reasoning spread fast. The B08 chapter tracks the lags from o1-preview. DeepSeek's R1-Lite preview came in 69 days, Alibaba's QwQ in 77, Google's Gemini 2.0 Flash Thinking in 98, R1 and Moonshot's Kimi k1.5 in 130, xAI's Grok 3 Think in 158, and Anthropic's Claude 3.7 Sonnet with visible extended thinking in 165. Google's Gemini 2.5 Pro (March 2025) took the top of public leaderboards as a thinking model by default. Meta's first flagship reasoning model came much later. The pattern is that an existence proof plus a published recipe compresses diffusion from years to weeks. See the Spread page for the full chart.

Deep Research and Claude Code

OpenAI's Deep Research (February 2025), an o3 variant trained with RL to browse and write cited reports, was among the first widely used agents that reliably finished long tasks. On 24 February Anthropic released Claude Code, an agent that works in a developer's terminal, alongside Claude 3.7 Sonnet. Andrej Karpathy coined "vibe coding" for writing software by describing it. Reasoning, tool use and verifiable tasks together had produced models that could take actions. See B10 for a deep dive.

2025

Coding agents, Meta's hiring spree, GPT-5 and Chinese open models

Every major lab shipped a coding agent, Chinese open-weights models closed the gap, and Meta tried to buy its way back to the frontier with a $14 billion Scale AI stake and nine-figure offers.

Coding agents from every lab

From April 2025 every lab shipped an agent for software. OpenAI's o3 and o4-mini used tools inside their reasoning, OpenAI's Codex cloud agent (May) ran tasks in parallel sandboxes, GitHub's Copilot coding agent turned issues into pull requests, and Google's Gemini CLI (June) offered a free terminal agent. Claude Opus 4 (May 2025) was Anthropic's first model deployed under its stricter ASL-3 safeguards. By autumn Claude Sonnet 4.5 was reported to stay on a single task for over 30 hours. Cursor became the fastest-growing developer tool, reaching a $29.3 billion valuation in November.

Benchmarks followed the shift from answering to doing. They included SWE-bench Verified for real GitHub issues, Terminal-Bench, OSWorld for using a computer, and METR's time-horizon measure, which found the length of tasks agents could complete doubling roughly every seven months. See B10 and B23 for deep dives.

Meta's Llama 4 miss and the hiring spree

Meta's Llama 4 (April 2025) disappointed, and its leaderboard results were criticized for using a specially tuned variant. Mark Zuckerberg responded by rebuilding. In June Meta took a 49% stake in Scale AI for about $14 billion and made its founder, Alexandr Wang, chief AI officer of a new Meta Superintelligence Labs. Nat Friedman joined. Daniel Gross left SSI for Meta after Sutskever rejected an acquisition, Shengjia Zhao left OpenAI to become MSL's chief scientist, and Ruoming Pang left Apple for a package reported in the hundreds of millions. Offers to individual researchers reportedly reached nine figures. OpenAI's leaders described it as someone breaking into their home.

In July, as Meta's hiring continued, Google paid $2.4 billion to license Windsurf's technology and hire its CEO after an OpenAI acquisition collapsed, and Cognition bought the rest of Windsurf days later. The pattern from 2012 had come back at a hundred times the price. Labs buy teams, often through licensing deals structured to avoid acquisition review. See Teams for the flow chart.

Olympiad gold and GPT-5

In July 2025 experimental models from OpenAI and Google DeepMind each reached gold-medal-level scores at the International Mathematical Olympiad, writing natural-language proofs under contest time limits. A year earlier the best result had been DeepMind's AlphaProof, a specialized system working in a formal proof language. Reasoning RL had generalized well beyond that formal setting.

GPT-5 (August 2025) unified OpenAI's fast and reasoning models behind a router that decides when to think. The launch was received as an incremental step, which revived the debate over whether progress had slowed. The same week OpenAI released gpt-oss, its first open-weights language models since GPT-2, and Google showed Genie 3, which generates explorable worlds in real time.

Chinese open models close in

Through 2025 Chinese labs released strong open-weights models at a pace US labs did not match. Alibaba's Qwen3 family (April) had switchable thinking, and Moonshot's Kimi K2 (July) was a one-trillion-parameter mixture-of-experts trained with a new optimizer (MuonClip). Zhipu's GLM-4.5, MiniMax's M1 and M2, and DeepSeek's V3.1 and V3.2, which introduced its own sparse attention, followed. Download counts for Chinese open models overtook American ones. Architecture innovation (linear and sparse attention hybrids for long context) increasingly came from these labs. See B20 and B02 for deep dives.

Veo 3, Nano Banana, Sora 2 and the lawsuits

Google's Veo 3 (May 2025) generated video with synchronized dialogue and sound in one pass. OpenAI's GPT-4o image generation (March) set off a trend of Studio Ghibli-style images, and Google's "Nano Banana" (August) made instruction-based photo editing consistent enough to be a mass product. ByteDance's Seedance 1.0 (June) joined Kling, Hailuo and Alibaba's open Wan as Chinese leaders in video. OpenAI's Sora 2 (September) launched as a social app with "cameos" of real people. Disney and Universal sued Midjourney in June, and music labels moved from suing Suno and Udio to licensing deals by year-end. See B11 and B13 for deep dives.

Sutskever's "age of research"

By late 2025 the most-quoted framing came from Sutskever in a November interview. He said the field was moving from an age of scaling back to an age of research, and that the central unsolved problem was that models generalize far worse than people and cannot learn continuously from experience. Gemini 3 (November 2025) took the top of most leaderboards and Google argued pretraining still had headroom, and Claude Opus 4.5 cut Anthropic's flagship price by two-thirds. Yann LeCun left Meta to found a company built on world models instead of language models. Two camps formed, one saying to keep scaling the current recipe and the other saying new ideas were needed, and each had evidence. See Next bets.

2026, first half

Anthropic gates Mythos and SpaceX agrees to buy Cursor

Anthropic gave its top model, Claude Mythos Preview, only to vetted partners for cybersecurity reasons and labs began splitting their models into capability tiers, while money and talent consolidated into a few very large groups. (From the atlas's release and talent logs, checked against sources.)

Claude Mythos Preview and Project Glasswing

The most consequential 2026 development in the atlas logs is Anthropic's April release of Claude Mythos Preview, a model above the Opus line that the company said found software vulnerabilities better than all but the most skilled humans. It went to a restricted partner programme, Project Glasswing, with $100 million in credits to find and fix vulnerabilities in critical software, and there was no broad launch. In June the public version, Claude Fable 5, shipped with classifiers routing cyber, biological and chemical requests, while Mythos 5 without those safeguards stayed with vetted partners. Three days later, per the logs, a US export-control directive forced Anthropic to suspend both for all users before access returned.

Even if some details change, the release pattern is clear. The strongest models are no longer released as one product to everyone. Labs ship capability tiers (OpenAI's GPT-5.x line split into Sol, Terra and Luna tiers by July), with the top tier governed like dual-use technology.

The Claude line after Opus 4.5

For your Opus question, the sequence is as follows. Claude Opus 4.6 (5 February 2026) added a one-million-token context and "agent teams" in Claude Code. Opus 4.7 followed on 16 April, and Claude Mythos Preview (7 April) opened the gated tier. Claude Opus 4.8 arrived on 28 May, only 42 days after 4.7, at the same $5/$25 price, with Anthropic reporting 69.2% on SWE-bench Pro against 64.3% for 4.7 and a model less likely to claim a task was finished when the evidence was thin. Then came Fable 5 and Mythos 5 (9 June), Claude Opus 5 (24 July, close to Fable 5 at half the price), Fable 5.1 (1 September), Claude Opus 5.5 (22 September, $4/$20) and Sonnet 5.5 (28 September). (The atlas's first release log missed Opus 4.8, and it is confirmed by Anthropic's own pricing page and launch coverage.) Anthropic also moved beyond coding. Cowork (January) brought the Claude Code agent to non-programming work, and Managed Agents (April) offered hosted agent infrastructure.

OpenAI's GPT-5.3-Codex, GPT-5.4 and the end of Sora

OpenAI's GPT-5.3-Codex (February) was described as the company's first model that was "instrumental in creating itself," with early versions helping debug its own training. That is the clearest public marker so far of the recursive loop the B24 chapter covers. GPT-5.4 (March) was OpenAI's first general model with native computer use, and GPT-5.5 followed in April. In March OpenAI announced it would shut down Sora, and the app went dark in April. Video was costly and contested, and the company was concentrating compute on agents and reasoning.

Chinese releases in 2026

ByteDance's Seedance 2.0 (February) generated audio and video jointly from many reference images, clips and sounds, and drew cease-and-desist letters from Disney, Netflix, Paramount, Warner Bros. and Sony within days. Zhipu's GLM-5 (February, MIT licence) and GLM-5.1 (April, tops SWE-Bench Pro, works up to eight hours on a task), Alibaba's Qwen3.5 (February) and DeepSeek-V4 (April, a 1.6-trillion-parameter mixture-of-experts with one-million-token context at a fraction of V3.2's compute) kept open models within months of the closed frontier. Meituan's LongCat-2.0 (June) was reported trained entirely on Chinese accelerators.

Talent moves and the Cursor sale

The talent flows in 2026 ran toward the biggest balance sheets. Barret Zoph and Luke Metz left Thinking Machines for OpenAI in January. Yann LeCun's AMI Labs raised a $1.03 billion seed in March to build JEPA-style world models. David Silver, AlphaGo's lead researcher, left Google DeepMind and founded Ineffable Intelligence, which raised $1.1 billion in April to pursue superintelligence through reinforcement learning on experience, without relying on human text. Andrej Karpathy joined Anthropic's pretraining team in May. Noam Shazeer, back at Google since 2024, moved to OpenAI in June. On the capital side, SpaceX agreed in June to buy Cursor's parent Anysphere for $60 billion, which put the leading coding tool inside Musk's AI unit.

2026, second half

GPT-6, Gemini 4 Argon and the labs betting on new research

As of October 2026 the newest frontier models, such as Gemini 4 Argon, reach vetted users first, and several well-funded labs, including SSI, AMI Labs and Ineffable Intelligence, are betting that the current recipe is not enough. (From the atlas's logs, checked against sources.)

July to September releases

July brought GPT-5.6's three tiers, Moonshot's 2.8-trillion-parameter open Kimi K3, Claude Opus 5 and ByteDance's Seedance 2.5 (30-second clips with synced audio). In August Alibaba released open weights for a 2.4-trillion-parameter Qwen3.8, and Meta, after a year closed, released Muse Glimmer 30B, an open model distilled from its proprietary Muse Spark. September was the densest month in the logs, with Anthropic's Fable 5.1, OpenAI's GPT-6 Astra (the company's first model rated Critical for cybersecurity in its own framework), then GPT-6 Sol and Luna, Claude Opus 5.5, DeepSeek-V4.1-Flash, and on 30 September Google DeepMind's Gemini 4 Argon, released first to vetted users. Gated-first release has become the norm for frontier-tier models.

Hassabis steps back at Google DeepMind

In August Demis Hassabis handed day-to-day running of Google DeepMind to Koray Kavukcuoglu and became Alphabet's chief scientist and DeepMind chair. That ends an arc that started with the 2014 acquisition. The founder who wanted independence to pursue science now sits at the top of Alphabet's research, while operations run as a product division. World Labs, Fei-Fei Li's spatial-intelligence company, agreed to join AMD in September, with Li becoming AMD's chief scientist.

SSI, AMI Labs and the other research bets

Two years in, SSI had still released no model, paper or product, by design. The logs record a July 2026 deal with Nvidia that gave it access to next-generation Vera Rubin systems for roughly ten times its previous compute, and Sutskever saying the lab had found research "worthy of scaling up." SSI, AMI Labs (world models), Ineffable Intelligence (RL from experience), Thinking Machines (customizable models and training tools) and Periodic Labs (AI with autonomous physical labs) are the clearest expressions of the "age of research" thesis. None had shipped a frontier model by October 2026. Whether one of them finds the missing ingredient, or the incumbents' scaled recipe keeps absorbing every idea first, is the main open question from here.

Recurring patterns, 2012 to 2026

Read back across fourteen years and five patterns repeat.

  1. Capability follows compute, but the binding constraint keeps moving. It was data in 2012, architecture in 2017, money in 2019, human feedback in 2022, inference cost in 2024, and power and chips in 2025 and 2026.
  2. Each breakthrough creates the next bottleneck. Scaling created the need for alignment; RLHF created sycophancy; reasoning created cost; agents created security risk.
  3. Diffusion keeps getting faster. Google kept the Transformer advantage for years; OpenAI kept reasoning for about two months; open-weights labs now ship within months of the frontier.
  4. People carry the ideas. The 2012 auction, the 2020 Anthropic split, the 2024 OpenAI departures and the 2025 poaching war each moved the frontier as much as any paper.
  5. Verifiable domains move first. Games came first, then code, then math, then agentic tasks with checkable outcomes. Wherever success can be scored automatically, RL finds a way to improve.