The week in AI
The week in brief
On 16 September Anthropic merged Claude Cowork and chat into a single Claude, adding Claude Docs and Claude Slides and moving Claude Design into conversations.
Anthropic also opened its Life Sciences Verification Program on 17 September, which gives vetted biology teams models with looser biology safeguards. Alibaba's Qwen team released Qwen3.8-Omni-Flash, an agent model with a 1M-token context, on 17 September and the open-weights Qwen-Image-2.1 on 20 September.
On 19 September DeepSeek posted a paper describing the sandbox platform behind its reinforcement learning work. The paper reports about 3 million sandboxes a day from one unit of the platform.
Anthropic folds Cowork into Claude and adds Docs and Slides
Since 16 September, Claude Cowork and Claude chat have been one product, and Claude Docs, Claude Slides and Claude Design are available in any conversation.
Users no longer choose between a chat mode and a Cowork mode. A quick question and a report due at noon go to the same Claude, and that Claude keeps working on a task after the user closes the laptop.
Claude Docs and Claude Slides are new. Claude Design already existed and now works inside conversations. The merged product is rolling out to Pro and Max subscribers, and artifacts, including the new document and slide formats, are on every plan, including Free.
The change is in the Claude app only, and the API is unaffected.
Anthropic loosens biology safeguards for vetted research teams
On 17 September Anthropic opened the Life Sciences Verification Program, a beta that gives verified institutions Mythos, Opus and Sonnet models with biology safeguards loosened.
The program fixes a problem with Fable 5. When a request touched virology, toxicology or molecular design, Fable 5 handed it to Opus 5 instead of answering itself, and that blocked legitimate drug-discovery work. Anthropic published a separate post on improving Fable 5's biology safeguards on the same day.
Verified teams can apply for one of two grants, "Standard Use" or "High-risk Use". A grant applies across Claude Science, claude.ai, Claude Code and the API. Anthropic says it built the program in partnership with the US government.
The beta is open to institutions only. Individual researchers on Pro and Max plans come later, and Anthropic has not given a date.
Alibaba's Qwen ships Omni-Flash agent model and open Qwen-Image-2.1
Alibaba's Qwen team released Qwen3.8-Omni-Flash on 17 September, a natively multimodal model built to plan and finish tasks across text, audio and video, with a 1M-token context.
Qwen3.8-Omni-Flash uses the Qwen3.8-Next architecture, a sparse mixture of experts (MoE) design in which each token passes through only a few of the model's expert sub-networks. Alibaba trained it on text, audio and video together from the start. The aim is to keep its text ability while carrying its agent skills over to audio and video. Alibaba targets video editing, translation and music-video generation, and the technical report also presents a framework called Qwen-MM-Plugins. The model is available through Alibaba's API only.
On 20 September the team released Qwen-Image-2.1 as open weights under a restricted license. Its generation component has 7B parameters, and one model now handles both text-to-image and editing, including native generation and editing of transparent images. The earlier Qwen image line used a 20B MMDiT model for generation and separate Edit models.
Qwen's GitHub page calls Qwen-Image-2.1 its most powerful open-source image generation model. The input has no independent benchmark results for either model.
DeepSeek paper reports 3 million agent sandboxes a day
According to a paper DeepSeek posted to arXiv on 19 September, one unit of its DeepSeek Elastic Compute (DSec) platform runs about 3 million sandboxes a day for agent training.
Agent training with reinforcement learning (RL) runs the model through many attempts at a task, and each attempt needs an isolated environment where the model can run code or use tools. DSec is the production system DeepSeek says provides those environments.
The paper reports more than 380,000 sandboxes running at once and more than 5,000 created per second, with one unit spanning about 160 nodes. These are DeepSeek's own figures. DSec offers four kinds of sandbox (function calls, containers, microVMs and full virtual machines) behind one SDK, and it loads image layers on demand from 3FS, DeepSeek's distributed file system.
DeepSeek says it designed DSec together with its RL framework so that the state of each attempt is kept and reward hacking is reduced. The atlas has only partly confirmed the details of the paper.
Also in the news
- Gemini 3.8 Live Google released Gemini 3.8 Live and Live Extended Thinking on 15 September, audio-to-audio models for voice agents that can reason in the background mid-conversation, and Google lists 97.7% on Big Bench Audio.
- Qwen3.8-LiveTranslate On 18 September Alibaba released a simultaneous interpretation model that cuts average latency from 2.8 seconds to 2.3 seconds and names who is speaking.
- Bolt Forge On 14 September Bolt.new introduced Bolt Forge, an agent built on open-source models, with up to 50 times more usage at no extra cost until 14 October.
- Pika On 17 September Pika relaunched as a platform of more than 25 apps that route to its own models or to third-party models such as Seedance, GPT Image and MiniMax H3.
- Suno On 18 September Universal and Sony filed a second suit against Suno in Massachusetts, covering 60,202 recordings and alleging that Suno v6 was trained on outputs of earlier Suno models.