The week in AI

Week of 5 Jan to 11 Jan 2026

The week in brief

Anthropic published a second generation of its Constitutional Classifiers on 9 January and reports that they cut the extra compute cost of jailbreak defense from 23.7% to about 1%.

The week's main research result came from Anthropic's safety work on jailbreak defense. The rest of the week was product launches. OpenAI opened ChatGPT Health on 7 January and Anthropic followed with Claude for Healthcare on 11 January, the same day Google announced an open standard for shopping through AI agents.

Anthropic cuts the cost of its jailbreak classifiers to 1%

Anthropic reports that its next-generation Constitutional Classifiers add about 1% compute to a model's serving cost, down from 23.7% for the first generation.

Constitutional Classifiers are separate models that screen what goes into and comes out of a Claude model. They are trained on a written set of rules (the "constitution") that says which content is allowed and which is not. Their main target is the universal jailbreak, a single prompting technique that gets a model past its safeguards on many different harmful questions at once.

The first generation worked but was expensive. Anthropic reported that it cut jailbreak success from 86% to 4.4%, at the cost of 23.7% more compute and a 0.38% rise in refusals of harmless requests. That overhead made it hard to justify running the classifiers on every request.

The paper published on 9 January describes a redesign that Anthropic says keeps the same robustness against universal jailbreaks while costing about 1% extra compute. Anthropic also reports fewer refusals on harmless queries, but the summary gives no figure for the new refusal rate. All of these numbers are Anthropic's own measurements, and the work is a research paper with no product release attached.

Also in the news

  • ChatGPT Health launched on 7 January as a separate space in ChatGPT that connects to a user's medical records and wellness apps so that answers draw on their own health data.
  • Qwen3-VL-Embedding and Qwen3-VL-Reranker came out on 8 January from Alibaba's Qwen team as open models in 2B and 8B sizes for searching across text, images, screenshots and video in more than 30 languages.
  • Scribe v2 is the new ElevenLabs speech-to-text model, released on 9 January. ElevenLabs says it has the lowest word error rate on the FLEURS benchmark and can transcribe several languages within one file.
  • Claude for Healthcare launched on 11 January with HIPAA-ready tools, connectors to the CMS Coverage Database, ICD-10 codes and the NPI provider registry, and additions for life-sciences work.
  • Universal Commerce Protocol was announced by Google at NRF 2026 on 11 January. It is an open standard that covers an agent's purchase from product discovery to payment and allows checkout inside AI Mode and the Gemini app.