The week in AI

Week of 23 Mar to 29 Mar 2026

The week in brief

OpenAI said on 24 March it will shut down its Sora video app and the Sora video API, about six months after the app launched.

Anthropic gave Claude two new kinds of autonomy. On 23 March computer use came to the Claude desktop app, and on 24 March Claude Code got an auto mode where classifiers approve routine actions. Speech models also had a busy week. Mistral released Voxtral TTS on 23 March, and on 26 March Alibaba released Qwen3.5-Omni while Google and Cohere shipped voice products. At xAI, two more co-founders left.

OpenAI is shutting down the Sora app and video API

OpenAI will close the Sora app on 26 April and retire the Sora video API on 24 September 2026, with no replacement.

OpenAI announced the shutdown on 24 March. The app and the web version close on 26 April. The Videos API, which serves the sora-2 and sora-2-pro models, is retired on 24 September, and OpenAI's deprecation notice lists no successor model.

OpenAI said its compute is better spent elsewhere. It also said the Sora team will now work on world simulation research for robotics. Usage had been falling. Engadget reported that Sora app downloads dropped 32% from November to December 2025.

The shutdown also ends a partnership with Disney that OpenAI had planned around Sora.

Anthropic lets classifiers approve Claude Code actions in auto mode

Anthropic launched auto mode for Claude Code on 24 March and reports that its safety classifiers wrongly block 0.4% of actions on real traffic.

Before auto mode, Claude Code users could approve every action by hand or turn the checks off. Auto mode sits between the two. A model decides whether each action is safe enough to run without asking. It launched as a research preview for Team plans.

Two checks do the screening. A server-side probe reads tool outputs and looks for prompt injection, which is text planted in a file or web page that tries to give the agent new instructions. A transcript classifier running on Sonnet 4.6 then judges each proposed action in two stages, and it is reasoning-blind, so it sees the actions and the conversation but not Claude's own reasoning.

Anthropic published its error rates. On 10,000 samples of real traffic the full pipeline blocked 0.4% of actions that should have been allowed. The miss rate is higher. On 52 real cases where Claude was overeager and took an action the user didn't want, the classifier let 17% through. On 1,000 synthetic data exfiltration cases it missed 5.7%.

The day before, on 23 March, Anthropic opened computer use as a research preview for Claude Pro and Max subscribers in the desktop app. Claude can now open apps, drive the browser and run tools on the user's own machine. Anthropic first offered computer use through its API in 2024. In the app Claude tries connectors first and falls back to the screen. The same release improved Dispatch, which lets a user assign Claude tasks remotely.

Mistral releases Voxtral TTS, Alibaba releases Qwen3.5-Omni

Mistral released Voxtral TTS on 23 March, a 4 billion parameter open-weights speech model that clones a voice from three seconds of audio.

Voxtral TTS is Mistral's first text-to-speech model. It covers nine languages and can carry a cloned voice across languages without extra training. Mistral reports model latency of about 70 milliseconds on typical inputs and a real-time factor of about 9.7, which means it generates audio roughly ten times faster than the audio plays.

The API price is $0.016 per 1,000 characters. The weights use the CC BY-NC 4.0 licence, so commercial use requires the paid API. Mistral reports that human raters preferred it over ElevenLabs Flash v2.5, with a 68.4% win rate on voice customization.

Alibaba's Qwen team released Qwen3.5-Omni on 26 March in Plus, Flash and Realtime versions, with no weights. It has two parts. The Thinker understands the input, and the Talker produces speech. Both use a mixture-of-experts design, where only some of the model's sub-networks run for each token. Alibaba's technical report describes hundreds of billions of parameters, a 256K token context and training on more than 100 million hours of audio-visual data.

Alibaba reports leading results on 215 audio and audio-visual subtasks for the Plus version. It says the model beats Gemini 3.1 Pro on key audio tasks and matches it on audio-visual understanding. The model takes more than ten hours of audio or 400 seconds of 720p video. A new method called ARIA aligns text units with speech units so streamed speech comes out steadier.

Two more voice releases came on 26 March. Google released Gemini 3.1 Flash Live, an audio-to-audio model for real-time voice agents, and put it into Search Live and Gemini Live in more than 200 countries. Cohere released Cohere Transcribe, an open speech recognition model that Cohere calls a new state of the art for open source.

Also in the news

  • Suno released Suno v5.5 on 26 March, which lets Pro and Premier users train up to three custom models on their own tracks, sing with their own voice through Voices, and use a free personalisation feature called My Taste.

People

  • Manuel Kroiss, an xAI co-founder who led pretraining and reported directly to Elon Musk, told colleagues the week of 23 March that he was leaving.
  • Ross Nordeen, an xAI co-founder who came from Tesla Autopilot, left on 27 March. Business Insider reported that he was the last of xAI's 11 co-founders other than Musk to leave, and Musk said xAI would be rebuilt from the foundations.