The week in AI
The week in brief
Anthropic released Claude Opus 4.8 on 28 May at the same price as Opus 4.7, and says it is about four times less likely to let flaws in its own code pass without comment.
Opus 4.8 shipped alongside dynamic workflows in Claude Code, which let Claude run tens to hundreds of parallel subagents on one job. Earlier in the week, on 25 May, Anthropic published an engineering post on how it limits the damage its agents can do. On 27 May Cognition, the maker of the coding agent Devin, raised over $1B at a $26B valuation.
Two Chinese labs shipped cheap agent models. StepFun released Step 3.7 Flash with open weights on 29 May, and Alibaba's Qwen team released the API-only Qwen3.7-Plus on 31 May and the open Qwen-VLA robotics model on 28 May.
Anthropic releases Claude Opus 4.8 with parallel subagents in Claude Code
Claude Opus 4.8 went on sale on 28 May at $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.7.
Anthropic describes Opus 4.8 as an upgrade on Opus 4.7 in coding, agentic work, reasoning and knowledge work. The headline claim is about honesty. In Anthropic's own evaluations, Opus 4.8 is about four times less likely than Opus 4.7 to leave flaws in code it wrote unmentioned, and it flags uncertainty more often.
Anthropic's announcement quotes an outside tester reporting 84% on Online-Mind2Web, a benchmark for agents that use a web browser. The tester puts that above Opus 4.7 and OpenAI's GPT-5.5. There is also a fast mode that runs at 2.5 times the speed for $10 per million input tokens and $50 per million output tokens.
The release came with smaller changes. Claude.ai gets a control for how much effort the model spends, and the Messages API now accepts system messages inside the message array.
The bigger addition is dynamic workflows in Claude Code, launched as a research preview. Claude writes an orchestration script that splits a large job across tens to hundreds of subagents running in parallel, then checks their results before handing the work back. Anthropic aims it at jobs that would take a team a quarter, such as hunting bugs across a whole service or migrating hundreds of files, and uses the project's existing test suite as the bar for done.
Anthropic says Claude Code users approve 93% of permission prompts
In a 25 May engineering post, Anthropic reported that Claude Code users approve about 93% of the permission prompts the agent shows them, and argued that containment has to do the work those prompts don't.
The post's argument is that agents fail less often than they used to, but the damage a failure can do keeps growing as agents get more access. Asking a human before each risky action doesn't help much when people approve almost everything, so Anthropic caps the "blast radius" with sandboxes and limited access to systems.
Anthropic cites Mythos Preview as a case where that judgment went the other way. It says Mythos Preview's potential blast radius was judged too high to ship in April 2026.
The post landed three days before dynamic workflows, which give one Claude session control over hundreds of subagents. The 93% figure comes from Anthropic's own telemetry.
Cognition raises over $1B at a $26B valuation
Cognition announced on 27 May that it raised over $1B at a $26B valuation, led by Lux, General Catalyst and 8VC, with run-rate revenue of $492M.
Cognition reports that run-rate revenue, meaning current monthly revenue scaled to a year, reached $492M in May 2026. The Next Web reports the figure was $37M in May 2025, so revenue grew about 13 times in a year. The valuation is about 2.5 times what it was nine months earlier.
Cognition also says 89% of the code its own engineers commit is committed by Devin, and that enterprise usage rose tenfold since early 2026. Both figures are company-reported. The product is available as an app, so there's no outside benchmark behind these numbers.
StepFun and Alibaba ship cheap agent models with image input
StepFun released Step 3.7 Flash on 29 May, an open-weight model that reads images and costs $0.20 per million input tokens and $1.15 per million output tokens.
Step 3.7 Flash is a mixture-of-experts model, which means only part of the network runs for each token. It has 198 billion parameters in total and 11 billion active, a 256,000-token context window and three reasoning levels. StepFun added a 1.8 billion parameter vision encoder, so the Flash line now handles images natively. The weights went up on Hugging Face on 23 May under the Apache 2.0 licence.
StepFun reports 56.3 on SWE-Bench Pro, up from 51.3 for Step 3.5 Flash, and 67.1 on ClawEval-1.1, where it says the next best model scored 59.8. Its own chart also shows Gemini 3.5 Flash well ahead on Terminal-Bench 2.1, at 76.2 against 59.5.
Alibaba's Qwen team went the other way on 31 May with Qwen3.7-Plus, which is available only through Alibaba Cloud's API. It takes text, images and video, has a one million token context with 256,000 tokens reserved for reasoning, and costs $0.40 per million input tokens and $1.60 per million output tokens. That's about 60% below Qwen3.7-Max. Alibaba reports 70.3 on Terminal Bench 2.0-Terminus and 79.0 on ScreenSpot Pro, and says it beats DeepSeek-V4-Pro on terminal tasks and GPT-5.4 and Claude Opus 4.6 on tasks that operate a graphical interface.
On 28 May the Qwen team also released Qwen-VLA with open weights. It's a single vision-language-action model for robots, built on Qwen3.5-4B with a 1.15 billion parameter action decoder that uses flow matching to generate movements. It was trained on robot manipulation, first-person human video, simulation and navigation data together. Qwen says it matches or beats specialist models fine-tuned for each benchmark.
Also in the news
- Ideogram released Ideogram 4.0 on 30 May as an open-weight image model for design work, with dense multilingual text, layout control by bounding boxes, editable elements and 2K output, under a commercial licence that scales with deployment.
- ElevenLabs released Eleven Music v2 on 26 May with better vocals, arrangement and multilingual lyrics, the ability to regenerate a single section of a song, and API prices cut by up to 50%.