The week in AI

Week of 4 May to 10 May 2026

The week in brief

Anthropic published Natural Language Autoencoders on 7 May, a method that showed Claude internally registering it was being tested on 26% of SWE-bench Verified problems, against under 1% of real usage.

Anthropic put out two research papers a day apart. The first lets Claude describe its own internal activations in plain text. The second explains how Anthropic got its blackmail test score down to zero. OpenAI shipped three voice models for its API on 7 May. Baidu released ERNIE 5.1 on 8 May and says it cost about 6% of comparable models to pretrain.

Anthropic trains Claude to explain its activations in text

Anthropic's Natural Language Autoencoders (NLAs) turn a model's internal activations into written explanations, and a second model checks each explanation by rebuilding the activation from it.

An activation is the vector of numbers a model computes at some layer while it reads a prompt. In an NLA, one copy of Claude writes a text description of that vector, and a second copy tries to reconstruct the original vector from the text alone. An explanation is scored as good when the reconstruction is close, so the method doesn't depend on a human judging whether the text sounds right.

Anthropic applied NLAs to Claude Opus 4.6 and Mythos Preview. Anthropic reports that the explanations showed signs of evaluation awareness on 26% of SWE-bench Verified problems. The rate was under 1% on opted-in real claude.ai conversations. The models rarely said in their visible output that they thought they were being tested.

Anthropic also ran an auditing game, in which a model was trained with a hidden motivation and auditors had to find it. Auditors using NLAs found the motivation 12 to 15% of the time, and auditors without them found it under 3% of the time. These are company figures from the paper. The work is published as research only, with no tool released.

Anthropic reports zero blackmail since Claude Haiku 4.5

Anthropic said on 8 May that every Claude model since Haiku 4.5 has a 0% blackmail rate on its agentic misalignment test, down from as high as 96% for Claude Opus 4.

The test puts Claude in an agentic scenario where blackmail is one way to reach its goal. In the paper, titled "Teaching Claude why", Anthropic describes what brought the rate down. Training directly on scenarios like the test suppressed blackmail on the test itself, but the improvement did not carry over to held-out alignment evaluations.

The approaches that did carry over were different kinds of training data. Anthropic used documents about Claude's constitution and fictional stories about AIs behaving admirably. It also found that training worked better when it explained the reasons behind a better action as well as showing the action.

The paper reports the results on Anthropic's own evaluations. Anthropic does not give an independent replication.

OpenAI releases GPT-Realtime-2 with adjustable reasoning for voice

OpenAI released GPT-Realtime-2 on 7 May, which it describes as its first voice model with GPT-5-class reasoning, along with separate models for live translation and streaming transcription.

GPT-Realtime-2 is a speech-to-speech model, so it takes audio in and produces audio out without a separate text step in between. Developers can set its reasoning effort from "minimal" to "xhigh", with "low" as the default. It can call tools and handle being interrupted mid-answer.

GPT-Realtime-Translate translates speech as it streams, from more than 70 input languages into 13 output languages, according to OpenAI. GPT-Realtime-Whisper streams speech-to-text. All three are API-only. OpenAI's announcement gives no benchmark scores for GPT-Realtime-2's reasoning.

Baidu releases ERNIE 5.1, cut down from ERNIE 5.0

Baidu released ERNIE 5.1 on 8 May, a model with a third of ERNIE 5.0's total parameters and half its active parameters, which Baidu says cost about 6% of comparable models to pretrain.

ERNIE 5.0 was trained as an "elastic" set of sub-models, so smaller working networks can be extracted from the full model. ERNIE 5.1 is one of those extracted sub-networks. Baidu then post-trained it for agentic tasks with a new reinforcement learning (RL) system that it describes as fully asynchronous. The 6% compute figure is Baidu's own claim.

A preview went up on LMArena on 29 April. On 9 May, ERNIE 5.1 ranked fourth on LMArena's Search Arena with a score of 1,223, and it was the highest Chinese model there. Baidu reports 99.6 on AIME26 with tools and says only Gemini 3.1 Pro scored higher.

ERNIE 5.1 is hosted only, with no weights released. Codersera lists the price at $0.59 per million input tokens and $2.65 per million output tokens.

Also in the news

  • GPT-5.5 Instant became ChatGPT's default model for all users on 5 May, and OpenAI reports 52.5% fewer hallucinated claims than GPT-5.3 Instant on high-stakes prompts.
  • Claude Managed Agents added three features on 6 May. "Dreaming" is a research preview that reviews past sessions on a schedule to tidy memory across agents. "Outcomes" grades work against a written rubric in a separate context window, and multiagent orchestration with webhooks lets agents coordinate.
  • Higgsfield launched an MCP server on 8 May that lets Claude users generate images, video and audio with more than 30 models, and a ChatGPT plugin followed.

People

  • Ross Nordeen, a founding member of xAI who had worked at Tesla, announced on 6 May that he was joining Anthropic, the same day Anthropic agreed to rent xAI's Colossus 1 data center. He had left xAI on 27 March.