The week in AI

Week of 9 Feb to 15 Feb 2026

The week in brief

Google released an upgraded Gemini 3 Deep Think on 12 February, and the ARC Prize Foundation verified its score of 84.6% on ARC-AGI-2.

Google DeepMind also described Aletheia, a math research agent built on Deep Think, and Zhipu AI released GLM-5, a 744-billion-parameter open-weight model under the MIT license. Chinese labs released several other models during the week. MiniMax shipped M2.5 for coding, Ant Group shipped a trillion-parameter reasoning model, and ByteDance launched its Seed 2.0 family and, reportedly, the Seedance 2.0 video model.

OpenAI released a fast coding model served on Cerebras hardware. It also hired Peter Steinberger, the creator of OpenClaw, on 14 and 15 February. Two xAI co-founders left on consecutive days.

Google upgrades Gemini 3 Deep Think and shows Aletheia

Gemini 3 Deep Think, now built on Gemini 3.1 Pro, scored 84.6% on ARC-AGI-2, a result the ARC Prize Foundation verified.

Google says it retuned the model with scientists and engineers so it handles messy research problems as well as contest math. It also reports 48.4% on Humanity's Last Exam without tools and a Codeforces Elo of 3455. Those two figures are company-reported. Google AI Ultra subscribers get the model in the Gemini app, and researchers and enterprises get early API access.

On 11 February Google DeepMind described Aletheia, a math agent that runs on Deep Think. It works in a loop. It generates a proof, a natural-language verifier checks it, and the agent revises or admits it failed. It also uses search so it doesn't cite papers that don't exist.

DeepMind says Aletheia wrote a research paper with no human input. In a semi-autonomous sweep of 700 open problems from the Erdős database, it solved four on its own. DeepMind also reports that Deep Think reached up to 90% on IMO-ProofBench Advanced as it was given more inference compute, with human experts grading the proofs.

Zhipu AI releases GLM-5 as a 744B open-weight model

Zhipu AI released GLM-5 on 11 February under the MIT license, with 744 billion total parameters and 40 billion active per token.

GLM-5 is a mixture-of-experts model. Each token is routed to 8 of 256 expert sub-networks, so only a fraction of the weights run at a time. Zhipu trained it on 28.5 trillion tokens, up from 23 trillion, and the model is about twice the size of GLM-4.5. It uses DeepSeek Sparse Attention, which DeepSeek introduced to cut the cost of long contexts.

Zhipu reports 77.8% on SWE-bench Verified and 50.4% on Humanity's Last Exam with tools. Artificial Analysis gave it 50 on its Intelligence Index v4.0, the highest of any open-weight model at launch. The API costs $1.00 per million input tokens and $3.20 per million output tokens.

Zhipu trained GLM-5 for long agent tasks with slime, its asynchronous reinforcement learning (RL) system. It tuned inference for Chinese chips from Huawei, Moore Threads and Cambricon. The launch came about a month after Zhipu listed on the Hong Kong stock exchange on 8 January.

MiniMax and Ant Group ship cheaper coding and reasoning models

MiniMax released M2.5 on 12 February, a 229B mixture-of-experts model that MiniMax reports at 80.2% on SWE-bench Verified.

MiniMax also reports 76.3% on BrowseComp with context management and 51.3% on Multi-SWE-Bench. It says M2.5 finishes a SWE-bench Verified run 37% faster than M2.1. The Lightning variant runs at 100 tokens per second and costs $0.30 per million input tokens and $2.40 per million output tokens. MiniMax says an hour of output at that speed costs about $1, and the standard variant is half the Lightning price. The weights are public under a restricted license.

Ant Group's inclusionAI team released Ring-2.5-1T on 10 February under the MIT license. Ant calls it the first open trillion-parameter thinking model built on hybrid linear attention. Linear attention keeps a fixed-size summary of earlier tokens instead of comparing every token with every other one. Ring mixes one standard attention layer for every seven Lightning Linear layers.

Ant reports that past 32,000 tokens this cuts memory access by more than ten times and raises generation throughput by more than three times. It also reports gold-level results on IMO 2025 and CMO 2025, which it tested itself.

ByteDance launches Seed 2.0 and Seedance 2.0

ByteDance released its Seed 2.0 models in Pro, Lite, Mini and Code versions on 14 February, and says the Pro model beats GPT-5.2 on SuperGPQA with 68.7%.

Seed 2.0 is the new flagship generation behind ByteDance's Doubao assistant, and it is tuned for production traffic and long-tail knowledge. ByteDance reports that the Pro model scores 77.3% on BrowseComp, 46.9% on SWE-Bench Pro and 55.8% on Terminal Bench 2.0. It also claims gold-level results on IMO, CMO and ICPC, with token prices about ten times lower than rival models. The models are available through a closed API.

On 12 February ByteDance reportedly launched Seedance 2.0, which generates audio and video together in a single model. A prompt can include up to nine images, three audio clips and three video clips as references, and output clips run up to 15 seconds. The atlas has not confirmed the details that followed. Reportedly, viral clips of copyrighted characters led Hollywood studios to send cease-and-desist letters within days, and ByteDance delayed the API rollout.

Also in the news

  • GPT-5.3-Codex-Spark went to ChatGPT Pro users as a research preview on 12 February. It is a smaller GPT-5.3-Codex served on Cerebras hardware, it is text-only with a 128K context, and OpenAI claims about 1,000 tokens per second. It is the first product of the OpenAI and Cerebras partnership announced on 14 January.
  • OpenClaw, the most-starred personal-agent project, is moving into an independent 501(c)(3) foundation that OpenAI supports, Sam Altman said on 14 February. Amazon, OpenAI and Red Hat are among its donors.
  • Gluon amplitudes are the subject of a preprint OpenAI posted on 13 February with physicists from IAS, Vanderbilt, Cambridge and Harvard. GPT-5.2 proposed a formula showing that a "single-minus" gluon tree amplitude is nonzero under specific kinematics, which many physicists expected to vanish. An internal OpenAI model proved the formula and the human authors verified it.
  • Qwen-Image-2.0 from Alibaba, released on 10 February, does image generation and editing in one model. It accepts instructions up to 1,000 tokens long for posters and slides, outputs native 2K images and is smaller than the 20B original. It is hosted on Qwen Chat and Alibaba Cloud, and the atlas found no open weights.

People

  • Mrinank Sharma left Anthropic, where he led safeguards research, on 9 February with a letter warning that "the world is in peril". Anthropic said he did not lead safety or its wider safeguards work.
  • Tony (Yuhuai) Wu, the xAI co-founder who led reasoning, announced his exit on 9 February, a week after SpaceX acquired xAI. Entrepreneur counted six of the 12 founders gone.
  • Jimmy Ba, an xAI co-founder and University of Toronto professor whom CNBC credited with research behind Grok 4, left on 10 February.
  • Ryan Beiermeister, OpenAI's top safety executive, was fired over alleged discrimination, the Wall Street Journal reported on 10 February. She called the allegation false and had opposed ChatGPT's "adult mode", and OpenAI said the firing was unrelated to any issue she raised.
  • Zoe Hitzig, an OpenAI researcher for two years, reportedly resigned on 11 February in a New York Times essay that objected to ChatGPT advertising and the risks of its archive of intimate user data.
  • Peter Steinberger, creator of OpenClaw, joined OpenAI on 15 February to work on next-generation personal agents. He reportedly turned down a competing offer from Meta.