The week in AI

Week of 16 Feb to 22 Feb 2026

The week in brief

Google released Gemini 3.1 Pro in preview on 19 February, and the ARC Prize verified its score of 77.1% on ARC-AGI-2, more than double the 31.1% of Gemini 3 Pro.

Three frontier labs shipped models within four days. Alibaba opened the weights of Qwen3.5-397B-A17B on 16 February. Anthropic made Claude Sonnet 4.6 its default model on 17 February, and Google followed with Gemini 3.1 Pro two days later.

On 20 February OpenAI published attempts by an internal model at the First Proof mathematics challenge, and Anthropic opened a research preview of Claude Code Security. World Labs announced $1 billion in new funding on 18 February.

Google's Gemini 3.1 Pro scores 77.1% on ARC-AGI-2

Gemini 3.1 Pro reached 77.1% on ARC-AGI-2 in a result verified by the ARC Prize, up from 31.1% for Gemini 3 Pro.

Google DeepMind released Gemini 3.1 Pro on 19 February as a preview aimed at agentic workflows. It is the upgraded core model behind Google's Deep Think update. ARC-AGI-2 is a set of visual reasoning puzzles built to be easy for people and hard for models, and the ARC Prize checked the score independently.

Google reports 80.6% on SWE-bench Verified with a single attempt and 94.3% on GPQA Diamond. Google also reports 68.5% on Terminal-Bench 2.0 using the Terminus-2 harness. On Humanity's Last Exam without tools, the company reports 44.4%. Google's comparison table puts Gemini 3.1 Pro next to Claude Opus 4.6 and GPT-5.2 and GPT-5.3-Codex, and every figure in that table is self-reported.

Google also shipped a separate endpoint called customtools. It steers the model toward a developer's own tools and away from defaulting to bash commands. That matters to teams that build agents on top of their own internal APIs.

Alibaba opens Qwen3.5, a 397B model with 17B active

Alibaba released Qwen3.5-397B-A17B under Apache 2.0 on 16 February, with native vision and 201 languages.

Qwen3.5-397B-A17B has 397 billion parameters and uses 17 billion of them for each token. It combines two techniques. A sparse mixture of experts routes each token to 11 of 512 expert sub-networks. Gated DeltaNet is a form of linear attention that keeps a fixed-size memory updated token by token, so each new token costs about the same no matter how long the context gets.

Alibaba says the model matches Qwen3-Max, its earlier flagship, while decoding 8.6 to 19 times faster. It was trained on text and images together from the start. Language coverage went from 119 languages in the previous generation to 201. The open weights support 262,000 tokens of context, and the hosted Qwen3.5-Plus offers 1 million.

On the Hugging Face model card, Alibaba reports 87.8 on MMLU-Pro and 88.4 on GPQA, and 85.0 on the multimodal benchmark MMMU. The coding figures differ between Alibaba's own sources. The model card lists 76.4 on SWE-bench Verified and Alibaba's blog lists 80.0 for the hosted setup.

The same day the Qwen team also released WebWorld, a family of open models in 8B, 14B and 32B sizes that simulate how websites respond to an agent's actions. Alibaba reports that training Qwen3-14B on trajectories generated in WebWorld raised its WebArena score by 9.2%.

Anthropic makes Claude Sonnet 4.6 its default model

Claude Sonnet 4.6 became the default model for Free and Pro users on 17 February, at $3 per million input tokens and $15 per million output tokens.

Claude Sonnet 4.6 keeps the price of Sonnet 4.5 and adds a 1 million token context window in beta. Anthropic describes it as an upgrade across coding, computer use, long-context reasoning, agent planning and design work.

In Anthropic's own testing in Claude Code, developers preferred Sonnet 4.6 over Sonnet 4.5 about 70% of the time. They preferred it over Opus 4.5, the larger model from the previous generation, 59% of the time. Anthropic reports 80.2% on SWE-bench Verified with a modified prompt and 60.4% on ARC-AGI-2 at high effort.

Three days later, on 20 February, Anthropic opened a limited research preview of Claude Code Security for Enterprise and Team customers. The tool reads a codebase and reasons about business logic and access control to find vulnerabilities. It checks each finding in several stages, and a human has to approve any patch it proposes. Anthropic says Claude Opus 4.6 found more than 500 vulnerabilities in production open-source code, some of which had gone undetected for decades. Open-source maintainers can apply for free expedited access.

OpenAI's internal model attempts all ten First Proof problems

OpenAI said on 20 February that an internal model attempted all ten First Proof problems and that at least five of its proofs are likely correct.

First Proof is a challenge set by leading mathematicians. It asks for complete, checkable proofs of research-level problems in specialist areas. OpenAI submitted attempts produced with limited human supervision.

OpenAI judges five of its proofs likely correct and names them as problems 4, 5, 6, 9 and 10. It also withdrew one claim. After the organizers published their official commentary, OpenAI conceded that its answer to problem 2 was wrong. The five likely correct proofs are OpenAI's own judgement, and the organizers had not published a full grading of them as of OpenAI's post.

Also in the news

  • ByteDance delayed the API launch of its Seedance 2.0 video model and promised stronger safeguards on 16 February. Disney, Netflix, Paramount, Warner Bros. and Sony had sent cease-and-desist letters over clips showing their characters.
  • xAI put Grok 4.20 into public beta on 17 February with a multi-agent mode that reportedly runs four agents sharing the same weights, which check each other's work before giving a single answer.
  • World Labs announced $1 billion in new funding on 18 February from partners including AMD, Autodesk, Emerson Collective, Fidelity, NVIDIA and Sea, and did not disclose its valuation.
  • Google added Lyria 3 to the Gemini app in beta on 18 February, so users can make 30-second music tracks with vocals and cover art from a text or image prompt.