The week in AI
The week in brief
Anthropic announced Claude Mythos Preview on 7 April and kept it from general release, giving access only to partner organizations in a cyber defence program, after reporting 93.9% on SWE-bench Verified.
Anthropic had the busiest week. Mythos Preview arrived with Project Glasswing, the access program that governs who can use it. Claude Managed Agents followed on 8 April, and smaller agent releases came on 9 April.
Outside Anthropic, Zhipu AI released the open-weights GLM-5.1, which it reports tops SWE-Bench Pro among models available to the public. Meta launched Muse Spark, the first model from Meta Superintelligence Labs, and did not release its weights.
Anthropic gated Claude Mythos Preview to Glasswing partners
On 7 April Anthropic released Claude Mythos Preview, a model above Opus, only to a closed group of organizations, because it finds software vulnerabilities better than all but the most skilled humans.
Mythos Preview is Anthropic's strongest model so far and sits a tier above Opus. Anthropic reports large gains over Claude Opus 4.6 on coding and security benchmarks. It scores 93.9% on SWE-bench Verified against 80.8%, and 77.8% on SWE-bench Pro against 53.4%. On Terminal-Bench 2.0 it moves from 65.4% to 82.0%. All of these are company-reported.
The security result explains the release decision. On CyberGym, which tests whether a model can reproduce known vulnerabilities, Anthropic reports 83.1% against 66.6% for Opus 4.6. Anthropic says the model out-finds all but the most skilled humans at locating software flaws, so it withheld general release over cyber risk.
Access runs through Project Glasswing, which puts defenders first. The 12 launch partners include AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks, alongside Anthropic, and Anthropic says more than 40 organizations have access in total. They are meant to use Mythos Preview to find and fix flaws in critical software before attackers can.
Anthropic is putting $100 million in usage credits into Glasswing. It is also giving $2.5 million to Alpha-Omega and OpenSSF and $1.5 million to the Apache Software Foundation. Once the credits run out, participants pay $25 per million input tokens and $125 per million output tokens. Anthropic has not said when or whether the model will be offered more widely.
Anthropic launched Claude Managed Agents, a hosted agent harness
On 8 April Anthropic released Claude Managed Agents, which runs the agent loop and its infrastructure for customers at standard token rates plus $0.08 per active session-hour.
Managed Agents moves the agent harness from the customer's servers to Anthropic's. Anthropic hosts the sandboxes, memory, scoped permissions and scheduling, supports long-running sessions, and shows traces in the Console. Multi-agent coordination and self-evaluation are in research preview.
Anthropic's engineering post describes how it is built. The model loop, which Anthropic calls the brain, runs separately from the sandboxes where tools execute, which it calls the hands, and both write to an append-only session log. A session can start without waiting for a sandbox to spin up. Anthropic reports that this cut median time to first token by about 60% and the 95th percentile by more than 90%.
The 9 April releases build on the same idea. The advisor tool lets a cheaper Sonnet or Haiku model doing the work consult Opus partway through a task. Anthropic says this gets close to Opus quality at close to Sonnet cost, with a one-line change to an API call.
Zhipu AI released GLM-5.1 with open weights
On 7 April Zhipu AI released GLM-5.1 under the MIT licence and reported a score of 58.4 on SWE-Bench Pro, ahead of GPT-5.4 and Claude Opus 4.6.
GLM-5.1 is a post-trained version of GLM-5 on the same 744-billion-parameter base. That base is a mixture-of-experts (MoE) model, which sends each token through a small subset of specialist sub-networks, so only part of the model runs at a time. The weights are on Hugging Face.
Zhipu reports 58.4 on SWE-Bench Pro, against 57.7 for GPT-5.4 and 57.3 for Claude Opus 4.6 as reported in the press. It reports 63.5 on Terminal-Bench 2.0 and 68.7 on CyberGym, and claims parity with Opus 4.6 across 12 benchmarks. These are Zhipu's own figures.
The model can work on a single task for up to eight hours on its own. Zhipu's demos include a 655-iteration run that built a Linux desktop and a 6.9 times throughput gain on a vector database.
Mythos Preview came out the same day. Anthropic's reported 77.8% on SWE-bench Pro is well above GLM-5.1, but Mythos is gated and GLM-5.1 can be downloaded by anyone.
Meta released Muse Spark without open weights
On 8 April Meta Superintelligence Labs released Muse Spark, a closed multimodal reasoning model, and did not publish its weights.
Muse Spark is the first model in Meta's Muse family and the first from Meta Superintelligence Labs. Meta says it rebuilt the architecture, data and infrastructure from scratch. Its Llama flagships came with open weights, and Muse Spark does not.
Meta says the model handles images and text natively and can reason over images step by step. It uses tools and has a Contemplating mode that runs several agents in parallel. According to Meta's post, Muse Spark scores 58% on Humanity's Last Exam and 38% on FrontierScience Research in Contemplating mode.
Meta also says Muse Spark matches Llama 4 Maverick with more than ten times less pretraining compute. The atlas has confirmed only part of this record against Meta's post, and none of these figures has been independently checked. Meta says larger Muse models are in development and has not given dates.
Also in the news
- HappyHorse-1.0 On 10 April Alibaba revealed that it built this joint audio and video model, which had topped the Artificial Analysis video arena under an anonymous name, with a text-to-video Elo of 1,333 without audio, according to a summary hosted by fal.
- Claude Cowork On 9 April Anthropic took Cowork out of research preview on macOS and Windows and added an analytics API, OpenTelemetry support and role-based access controls for Enterprise customers.