The week in AI

Week of 13 Apr to 19 Apr 2026

The week in brief

Anthropic released Claude Opus 4.7 on 16 April, and the company reports 64.3% on SWE-bench Pro, up from 53.4% for Opus 4.6, at an unchanged price.

Anthropic shipped Opus 4.7 with its cyber capability deliberately held below its restricted Mythos Preview model. Two days earlier it published a study in which Claude agents ran alignment experiments faster than human researchers. OpenAI put two specialised models behind identity checks, GPT-5.4-Cyber for security defenders and GPT-Rosalind for biology. Coding tools moved further toward running many agents in the background, with releases from OpenAI, Cognition and Anthropic.

OpenAI also lost two senior leaders on 17 April. Bill Peebles, head of Sora, and Kevin Weil, who ran OpenAI for Science, both left.

Anthropic released Claude Opus 4.7 with 64.3% on SWE-bench Pro

Anthropic reports that Claude Opus 4.7 scores 64.3% on SWE-bench Pro and 87.6% on SWE-bench Verified, at $5 per million input tokens and $25 per million output tokens.

Opus 4.7 went out on 16 April through Anthropic's API. In the vendor-reported table, its SWE-bench Pro score is ahead of GPT-5.4 at 57.7% and Gemini 3.1 Pro at 54.2%. On SWE-bench Verified it moved from 80.8% for Opus 4.6 to 87.6%, still below the 93.9% Anthropic reports for Mythos Preview.

Two outside evaluators published numbers through Anthropic's announcement. Cursor measured 70% on its CursorBench, up from 58% for Opus 4.6. XBOW measured 98.5% on its visual-acuity benchmark, up from 54.5%, alongside a new image limit of 2,576 pixels on the long side, about 3.75 megapixels. The model also gets file-system memory, so it can write notes to files and read them back in later sessions.

Pricing is the same as Opus 4.6, but a new tokenizer turns the same text into roughly 1.0 to 1.35 times as many tokens. For some text, that means the same request costs up to about a third more.

Anthropic says it kept Opus 4.7's cyber capability intentionally below Mythos Preview. Requests that look like attacks are blocked automatically, and security professionals can apply to a Cyber Verification Program for fewer restrictions.

Claude agents closed 0.97 of an alignment gap that humans closed 0.23 of

In an Anthropic Fellows study published 14 April, parallel Claude agents reached a performance gap recovered of 0.97 on a weak-to-strong supervision task, against 0.23 for human researchers.

The task is weak-to-strong generalization. A small model labels data and a bigger model is trained on those labels, and researchers measure how much of the bigger model's full ability survives the weak teaching. Performance gap recovered (PGR) is that measure, where 0 means no better than the weak teacher and 1 means as good as training on true labels.

Here the teacher was Qwen 1.5-0.5B-Chat and the student was Qwen 3-4B-Base. Human researchers worked on it for seven days and reached a PGR of 0.23. The Claude agents then worked for five more days and reached 0.97.

Anthropic reports the agents used 800 cumulative hours and cost about $18,000, or $22 per agent-hour. The study is a paper only, with no product attached, and Anthropic measured the result itself.

OpenAI released GPT-5.4-Cyber and GPT-Rosalind to vetted users

OpenAI released two specialised models behind trusted access in three days, GPT-5.4-Cyber for security on 14 April and GPT-Rosalind for biology on 16 April.

GPT-5.4-Cyber is GPT-5.4 fine-tuned to refuse less on defensive security work. It sits inside OpenAI's Trusted Access for Cyber program, which OpenAI says now covers thousands of verified defenders. Users pass identity verification, and the most permissive tier needs an extra application. OpenAI framed the program as preparation for more capable models.

Simon Willison, who writes a widely read blog on AI tools, read it as OpenAI's response to Anthropic restricting Claude Mythos. Two days later, Anthropic's own Opus 4.7 launch used a similar setup, with automatic blocking plus a verification program for security professionals.

GPT-Rosalind is a reasoning model for biology and drug discovery, and OpenAI calls it the first in a life-sciences series. It is tuned for tool use across chemistry, protein engineering and genomics, and it comes with a free Codex plugin that connects more than 50 scientific tools. OpenAI named Amgen, Moderna, Novo Nordisk and Thermo Fisher as customers, and access goes to qualified customers only.

OpenAI, Cognition and Anthropic built coding tools for background agents

On 16 April OpenAI gave Codex background computer use, an in-app browser and more than 90 plugins, a day after Cognition shipped Windsurf 2.0 with built-in Devin cloud agents.

Codex can now operate any Mac app with its own cursor while the user keeps working in parallel. It also gained image generation, memory, pull request review, and automations it can schedule for itself across several days. OpenAI titled the release "Codex for (almost) everything" and pitched it beyond programmers.

Windsurf 2.0 from Cognition, released 15 April, adds an Agent Command Center. That's a kanban-style board showing local and cloud agents side by side. Spaces carry project context between sessions, and developers can hand tasks to Devin from inside the editor. Devin Cloud is included in every Windsurf plan, with a gradual rollout. Cursor 3 made a similar agent-first change about two weeks earlier.

Anthropic's Routines in Claude Code, a research preview from 14 April, run on Anthropic's cloud. A routine is a prompt plus a repository and connectors, triggered by a schedule, an API call or an event, so it runs without the developer's laptop open. Each routine gets its own endpoint and token. Anthropic's example pulls the top Linear bug each night and opens a draft pull request. The earlier /schedule command became scheduled routines.

Also in the news

  • Claude Design is an Anthropic Labs research preview on Opus 4.7, released 17 April for paid plans, that makes designs, prototypes, slides and one-pagers with a team's design system applied and hands off to Claude Code.
  • Qwen3.6-35B-A3B and Qwen3.6-27B are open-weight Apache 2.0 coding models from Alibaba, released 16 April, which Alibaba reports at 73.4 and 77.2 on SWE-bench Verified, with a "preserve_thinking" option that keeps earlier reasoning in context across agent turns.
  • Qwen3.6-Max-Preview is the largest Qwen3.6 tier, previewed on 14 April through Alibaba Cloud only and aimed at front-end work.
  • Gemini Robotics-ER 1.6 from Google DeepMind, released 14 April, improves pointing, counting, success detection and multi-view spatial reasoning, and adds instrument reading developed with Boston Dynamics.

People

  • Bill Peebles left OpenAI on 17 April after three years leading Sora, as OpenAI shut down Sora, which TechCrunch reported cost about $1 million a day in compute.
  • Kevin Weil, OpenAI's former chief product officer and head of OpenAI for Science, left on 17 April after that team was folded into other research groups, and Business Insider reports he now runs a science startup seeking a valuation above $750 million.