Anthropic announces Claude Mythos Preview to a gated group under Project Glasswing

Anthropic announced Claude Mythos Preview with Project Glasswing, a gated research preview described as a general-purpose frontier model and "our most capable yet for coding and agentic tasks," which…

Date
7 April 2026
Who
Anthropic
Confidence
High on announcement; Medium on benchmark figures (secondary)
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): Anthropic · Confidence: High on announcement; Medium on benchmark figures (secondary) Anthropic announced Claude Mythos Preview with Project Glasswing, a gated research preview described as a general-purpose frontier model and "our most capable yet for coding and agentic tasks," which had identified thousands of zero-day vulnerabilities. The preview came with twelve launch organisations (Anthropic itself, AWS, Apple, Google, Microsoft, Nvidia and others), over 40 more organisations and $100M in usage credits (Anthropic).

Secondary reports quote 93.9% on SWE-bench Verified, 94.6% on GPQA Diamond and 97.6% on USAMO 2026 (llm-stats). On 2026-06-09 Anthropic released Claude Fable 5, the same underlying capability class as a general release with classifiers that route cyber, biology, chemistry and distillation-flagged requests to the weaker Opus 4.8, and Mythos 5 for vetted defenders; both have always-on adaptive thinking and 1M-token context, at $10 / $50 per million input / output tokens (The Hacker News, release notes).

One independent datapoint is the UK AI Security Institute's evaluation (2026-04-13), which found Mythos Preview the first model to solve its 32-step "The Last Ones" corporate-network range end to end (3 of 10 attempts; an average of 22 steps against 16 for Claude Opus 4.6, the runner-up) and the first to reach 73% on expert-level capture-the-flag tasks, while stressing that its ranges lack active defenders and that it could not say whether the model would succeed against well-defended systems (AISI).

Anthropic's own claims remain company-reported, and critics dispute parts of them. Wikipedia's summary (a pointer, not checked against the primaries here) records an independent researcher saying the zero-day counts cannot be verified outside promotional documents, a report that Opus 4.6 found some bugs that Mythos then exploited, and a drop in Mythos's full-code-execution rate to under 5% once the two most exploitable bugs were patched (Wikipedia; see the Backlog). For this chapter, the headline abilities are reasoning-driven (long-horizon agentic reasoning over code), and the release pattern shows that gating decisions now follow capability, whatever the reasoning style (B22, B10).

Read it in the deep dive