The week in AI

Week of 16 Mar to 22 Mar 2026

The week in brief

Cursor released Composer 2, its in-house coding model, on 19 March, and Cursor reports a score of 61.7 on Terminal-Bench 2.0 at $0.50 per million input tokens.

Coding models dominated the week. OpenAI shipped GPT-5.4 mini and GPT-5.4 nano on 17 March. Mistral released Mistral Small 4 and a Lean 4 proof agent called Leanstral on 16 March. MiniMax released M2.7 on 18 March and says the model helped improve itself during development.

The research side had two architecture papers on 16 March. One is Mamba-3 from CMU, Princeton and collaborators. The other is Attention Residuals from Moonshot AI's Kimi team.

Cursor releases Composer 2, built on Moonshot's Kimi K2.5

Cursor's Composer 2 scores 61.7 on Terminal-Bench 2.0, up from 47.9 for Composer 1.5, according to Cursor.

Anysphere, the company behind the Cursor editor, released Composer 2 on 19 March. Cursor took a base model and ran continued pretraining on it, which means more next-token training on new data. It then applied reinforcement learning (RL) on tasks that need hundreds of actions to finish, so the model learns to carry long coding jobs through to the end.

All of the scores below are Cursor's own. Composer 2 gets 73.7 on SWE-bench Multilingual, against 65.9 for Composer 1.5. On CursorBench, a benchmark Cursor built itself, it scores 61.3, against 44.2 for Composer 1.5.

The model runs only inside Cursor. It costs $0.50 per million input tokens and $2.50 per million output tokens. A fast tier costs $1.50 per million input tokens and $7.50 per million output tokens.

The base model is Moonshot AI's Kimi K2.5, licensed through Fireworks, and the launch post left this out. Cursor confirmed it in a later statement. So a leading American coding product now runs on top of an open Chinese model, with Cursor's training added on top.

OpenAI ships GPT-5.4 mini and nano at higher prices

OpenAI reports that GPT-5.4 mini scores 54.4% on SWE-Bench Pro, close to the full GPT-5.4's 57.7%.

OpenAI released GPT-5.4 mini and GPT-5.4 nano on 17 March for high-volume work. The company says mini runs more than twice as fast as GPT-5 mini. Nano is the cheapest model in the family and is meant for classification and for subagents, which are smaller models that a main agent hands tasks to.

The scores are OpenAI's own and use the highest reasoning setting, called xhigh. On SWE-Bench Pro, mini reaches 54.4% and nano 52.4%, against 45.7% for GPT-5 mini. On OSWorld-Verified, a test of operating a computer, mini scores 72.1% against 75.0% for GPT-5.4. Nano manages 39.0% there, so it falls well behind on computer use.

Prices went up. Mini costs $0.75 per million input tokens and $4.50 per million output tokens. Nano costs $0.20 per million input tokens and $1.25 per million output tokens. That is three times the input price of GPT-5 mini and four times that of GPT-5 nano.

Mistral releases Small 4 and the Leanstral proof agent

Mistral Small 4 is a 119 billion parameter model with 6 billion active, released on 16 March under Apache 2.0.

Mistral AI merged three earlier models into Mistral Small 4. Those are Magistral for reasoning, Pixtral for vision and Devstral for agentic coding. Small 4 is a mixture of experts (MoE), meaning each token goes to only a few of many sub-networks. It has 128 experts with 4 active per token, and the context window is 256K tokens. Users can set how much reasoning effort it spends.

Mistral claims 40% lower latency and three times the throughput of Small 3. It also claims parity with GPT-OSS 120B on shorter outputs. Both figures come from Mistral.

Leanstral uses the same sparse design, with 120 billion parameters and 6 billion active. It is an open agent for writing proofs in Lean 4, a language in which a computer checks every step of a proof. Mistral trained it so that people don't have to review each proof by hand.

On Mistral's own Lean evaluation, Leanstral scores 26.3 at pass@2 for $36. Pass@2 means the model gets two attempts per problem. Claude Sonnet scores 23.7 at a cost of $549. With sixteen attempts Claude Opus is still ahead, scoring 39.6 for $1,650 against Leanstral's 31.9 for $290.

A day later, on 17 March, Mistral launched Forge. Forge lets companies pretrain, post-train and RL-align models on their own data, and ASML, ESA and Ericsson are early partners.

MiniMax releases M2.7 and says it helped train itself

MiniMax says M2.7 ran more than 100 rounds of autonomous optimization, which gave a 30% gain on an internal version.

MiniMax released M2.7, a 229 billion parameter model, on 18 March and calls it its first self-evolving model. During development the model updated its own memory, built skills and evaluation sets, and changed its own scaffold, which is the code that runs the model as an agent. MiniMax describes this as recursive self-improvement. The figures given don't say which evaluation the 30% gain was measured on.

All the scores are MiniMax's own. M2.7 reaches 56.22% on SWE-Pro and 57.0% on Terminal Bench 2. It gets a 66.6% medal rate on MLE Bench Lite, which tests machine learning engineering tasks, and an Elo of 1495 on GDPval-AA.

The Terminal Bench 2 score is a little below Cursor's 61.7 for Composer 2. Each company ran its own evaluation, so the two numbers aren't a controlled comparison.

Also in the news

  • Mamba-3, from CMU, Princeton and collaborators, is a state space model designed for fast inference, with complex-valued state for tracking state and a multi-input multi-output update that adds quality without slowing decoding. The authors report 1.8 points over Gated DeltaNet with the MIMO update (0.6 without it) at 1.5 billion parameters, and the same perplexity as Mamba-2 with half the state size.
  • Attention Residuals, from Moonshot AI's Kimi team, replaces the fixed sum of earlier layers' outputs with learned attention weights over those outputs, and it showed gains in scaling runs and in Kimi Linear, a 48 billion parameter model with 3 billion active.
  • Xiaomi released MiMo-V2-Pro on 18 March, its first model of about a trillion parameters with 42 billion active, as a closed API model alongside V2-Omni and V2-TTS.
  • Midjourney reportedly opened an alpha of V8 on 17 March.
  • Google DeepMind proposed ten cognitive abilities as a framework for measuring general intelligence on 17 March, and opened a $200,000 Kaggle hackathon to build the evaluations.

People

  • Mustafa Suleyman, CEO of Microsoft AI, handed off Copilot on 17 March when Satya Nadella put it under Jacob Andreou, a Microsoft EVP who came from Snap, so Suleyman can focus on frontier models and the superintelligence effort while still reporting to Nadella.