The week in AI

Week of 20 Apr to 26 Apr 2026

The week in brief

OpenAI released GPT-5.5 on 23 April and reports 82.7% on Terminal-Bench 2.0, up from 75.1% for GPT-5.4.

DeepSeek followed on 24 April with open weights for DeepSeek-V4-Pro, a 1.6 trillion parameter model with a one million token context. Xiaomi, Moonshot AI and Tencent also released open-weight models in the same week. On 21 April SpaceX announced a partnership with Cursor that includes an option to buy the company for $60 billion.

Google spent the week on infrastructure. It announced eighth-generation TPUs and a training method that keeps working when whole groups of chips fail. OpenAI also shipped a new image model, gpt-image-2.

OpenAI releases GPT-5.5 with gains in agentic coding

**OpenAI** released GPT-5.5 and GPT-5.5 Pro on 23 April to paid ChatGPT and Codex users, and the API followed on 24 April with a one million token context window.

OpenAI reports 82.7% on Terminal-Bench 2.0, against 75.1% for GPT-5.4, and 78.7% on OSWorld-Verified, up from 75.0%. Terminal-Bench tests coding agents working in a command line. OSWorld tests whether a model can operate a desktop computer. On GDPval, OpenAI's comparison against professionals' work, the company reports wins or ties 84.9% of the time, a small step from 83.0%.

The largest company-reported jump is in mathematics. GPT-5.5 scores 35.4% on FrontierMath Tier 4, the hardest tier, compared with 27.1% for GPT-5.4. OpenAI says the gains come with fewer tokens spent per task, and it puts them in agentic coding, computer use and early scientific research.

The API arrived a day after ChatGPT. OpenAI says the delay was for safeguards. GPT-5.5 Pro is the variant for harder problems, and OpenAI has not published separate prices for it in the material behind this page.

DeepSeek opens V4 with a cheaper million-token context

**DeepSeek** released previews of DeepSeek-V4-Pro and V4-Flash on 24 April under the MIT licence, and reports that at one million tokens V4-Pro needs 27% of the compute and 10% of the KV cache of V3.2.

V4-Pro is a mixture-of-experts (MoE) model, which routes each token through a small subset of its weights. It has 1.6 trillion parameters in total and 49 billion active per token. V4-Flash has 284 billion in total and 13 billion active. Both were pretrained on more than 32 trillion tokens, and one million tokens is now the default context across DeepSeek's services.

The cost saving comes from attention. V4 mixes two kinds of compressed attention, which DeepSeek calls Compressed Sparse and Heavily Compressed, so the model stores and reads far less per token of context. The KV cache is the memory a model keeps for every earlier token, and it is usually what makes long contexts expensive to serve. DeepSeek also uses manifold-constrained hyper-connections, a change to how layers pass information forward, and trains with the Muon optimizer.

DeepSeek reports 80.6 on SWE-bench Verified for V4-Pro in its maximum reasoning setting and 79.0 for Flash. On GPQA Diamond the company reports 90.1 for Pro and 88.1 for Flash. Weights ship in FP4 and FP8, and the API names are deepseek-v4-pro and deepseek-v4-flash.

Around the same time, DeepSeek published TileKernels. It's an open library of kernels written in TileLang for MoE routing, FP8 and FP4 quantization and the hyper-connection layers that V4 uses.

SpaceX takes a $60 billion option on Cursor

**SpaceX** announced on 21 April that it is working with **Cursor** and has an option to buy the company for $60 billion later in 2026, or to pay $10 billion for the joint work instead.

The deal pairs Cursor's coding product and its users with SpaceX's Colossus compute. SpaceX says Colossus is roughly one million H100 equivalents. The two plan to train a model for coding and knowledge work.

Cursor had been seeking a funding round at a $50 billion valuation before the announcement. Two Cursor engineers had also moved to xAI before it. SpaceX will decide on the acquisition later in the year, and the terms beyond the two headline figures are not public.

Xiaomi, Moonshot and Tencent release open-weight models

**Xiaomi** released MiMo-V2.5-Pro on 22 April under the MIT licence, a model with 1.02 trillion parameters in total and 42 billion active and a one million token context.

Most layers in MiMo-V2.5-Pro use sliding-window attention, which only looks at nearby tokens, and one layer in seven looks at the whole context. Xiaomi says this cuts the KV cache by nearly seven times. The model has three multi-token prediction (MTP) layers, which predict several tokens ahead to speed up generation. It was pretrained on 27 trillion tokens in FP8 and then trained with agentic reinforcement learning (RL) and distillation from several teacher models.

On OpenRouter, V2.5-Pro costs $0.435 per million input tokens and $0.87 per million output tokens. The smaller MiMo-V2.5, with 310 billion parameters in total and 15 billion active, costs $0.14 per million input tokens and $0.28 per million output tokens.

Moonshot AI released Kimi K2.6 on 20 April on the same base as K2.5, with one trillion parameters in total and 32 billion active. Moonshot reports 80.2% on SWE-bench Verified and says its agent swarm now coordinates 300 sub-agents over 4,000 steps, up from 100 sub-agents and 1,500 steps. Moonshot also cites autonomous coding runs of 12 to 13 hours.

Tencent released Hy3 preview on 23 April. It is the first model from Tencent's rebuilt pretraining and RL stack under Yao Shunyu, Tencent's chief AI scientist. Hy3 is an MoE model with 295 billion parameters in total and 21 billion active, a 256,000 token context and a licence that restricts use. Tencent says it shipped the model early to collect product feedback.

Also in the news

  • gpt-image-2 OpenAI's image model, released on 21 April as ChatGPT Images 2.0, searches the web and checks its own output before drawing, outputs up to 2K, and renders small text and Japanese, Korean, Hindi and Bengali script much better than earlier models.
  • TPU 8t and TPU 8i Google split its eighth-generation TPU into a training chip and an inference chip on 22 April, and says a TPU 8t superpod scales to 9,600 chips and 121 exaflops, with general availability later in 2026.
  • Decoupled DiLoCo Google DeepMind described on 23 April a training method that runs separate groups of chips asynchronously with little communication between data centers, and in tests with Gemma 4 training continued after whole groups failed.
  • Deep Research Max Google released two Deep Research agents on Gemini 3.1 Pro in its API on 21 April, a fast one for live interfaces and a Max version for long overnight reports, both able to search the web, MCP servers and uploaded files.
  • Claude Code Anthropic said on 23 April that recent quality complaints came from three product changes, including lowering the default effort setting from high to medium, and that the model itself had not degraded.
  • FlashQLA Alibaba's Qwen team released TileLang kernels for Gated DeltaNet linear attention on 24 April, which it says run the forward pass two to three times faster than the existing FLA Triton kernels on Hopper and Blackwell GPUs.
  • Kling 3.0 Kuaishou announced on 23 April what it calls the first native 4K video generation, which avoids the artifacts that upscaling leaves in faces and fine detail.
  • Workspace agents OpenAI made its shared, Codex-powered agents generally available in ChatGPT Business, Enterprise and Edu on 22 April.