The week in AI
The week in brief
OpenAI released GPT-6 Astra on 3 September, its first model rated Critical for cybersecurity, and reports 97.6% on FrontierMath Tier 4.
Anthropic shipped Claude Fable 5.1 and its less restricted sibling Mythos 5.1 on 1 September, with a big jump on Terminal-Bench 4.0 and a 75% cut to the price of cache reads. On 4 September Anthropic reported that Claude agents had written a complete computer-checked proof of Fermat's Last Theorem in Lean.
OpenAI closed the week on 6 September by saying it had met its goal of an automated AI research intern. Google released Gemini 3.8 Flash and WeatherNext 3, Meta updated Muse Spark, and World Labs and Runway both showed new world models.
OpenAI released GPT-6 Astra, rated Critical for cybersecurity
OpenAI released GPT-6 Astra on 3 September at $10 per million input tokens and $50 per million output tokens, and it is the company's first model rated Critical for cybersecurity.
GPT-6 Astra is a closed API model with a phased rollout, and OpenAI gates access by trust level. It refuses advanced exploit work until a customer has Daybreak access. A Fast mode runs 2.5 times faster at twice the price.
OpenAI reports 97.6% on FrontierMath Tier 4 (v2), up from 83.0% for GPT-5.6 Sol. On ARC-AGI-3 it reports 99.9% using its own harness. Simon Willison reported 62.7% in the ARC default harness, so the headline figure depends heavily on the scaffolding around the model. OpenAI lists GPT-5.6 Sol at 7.8% on the same benchmark.
The coding and computer-use gains are smaller. OpenAI reports 57.9% on Terminal-Bench 4.0, against 37.3% for GPT-5.6 Sol, and 72.6% on OSWorld 2.0, against 65.7%. On the independent Artificial Analysis Intelligence Index v4.1.1, as listed in OpenAI's own table, Astra scores 61.2. That's barely above GPT-5.6 Sol at 60.9 and below Anthropic's Claude Fable 5.1 at 65.7.
The cyber rating follows OpenAI's tests without production safeguards. There, the model scored 100% on ExploitBench and 88.0% on SRE-Bench. OpenAI also says Astra's written reasoning is harder to monitor than GPT-5.6 Sol's. Press accounts say it was pretrained on more than 100,000 GPUs at Stargate Texas with a looped "recurrent depth" design, where the same layers are run several times over. It then had reinforcement learning (RL) for computer use, coding and science. OpenAI has not confirmed the architecture details in the records the atlas checked.
Anthropic shipped Claude Fable 5.1 and Mythos 5.1
Anthropic released Claude Fable 5.1 on 1 September, reporting 55.8% on Terminal-Bench 4.0, up from 42.0% for Fable 5.
Fable 5.1 is a same-tier upgrade of Fable 5. Mythos 5.1 is the same model with fewer safeguards, and Anthropic reports 60.9% for it on Terminal-Bench 4.0. The largest reported gain is on Terminal-Bench-Science 0.1, which rose from 24.7% to 52.6%. CursorBench 3.2.0 went from 70.5% to 73.4%.
The price stays at $10 per million input tokens and $50 per million output tokens. Cache reads drop from $1.00 to $0.25 per million tokens, and Anthropic says typical workloads come out about 25% cheaper.
Anthropic retuned the safeguards. It reports 60% fewer false positives on cyber requests, and the model may now find vulnerabilities but not exploit them. Anthropic also set up a biology access program with the US government and limited users' ability to edit the model's prior thinking, which it describes as an anti-distillation measure.
On the same day Anthropic announced Enterprise Frontier Safeguards, a response to customer objections over the 30-day data retention required on Mythos-class models. Misuse screening data stays in cloud infrastructure the customer controls. Anthropic built it with more than 100 customers and with AWS, Google Cloud and Azure, and it rolls out in phases from fall 2026. Eligible Fable users get zero data retention in the meantime.
Claude agents formalized Fermat's Last Theorem in Lean
Anthropic reported on 4 September that Claude agents wrote a complete Lean proof of Fermat's Last Theorem in about 11 days, totalling 13 million lines.
Lean is a proof assistant, a program that checks every step of a mathematical proof mechanically. Anthropic describes this as the first complete computer-checked proof of the theorem. The agents followed a simplified version of Andrew Wiles's proof.
Dozens of agents worked on it, coordinated through Columbia's Prove2Me platform and running an internal model Anthropic calls roughly comparable to Fable 5.1. They proved 30,300 theorems, and 29,500 of them are used in the final proof. Anthropic says that is over five times the size of Mathlib, Lean's main community maths library, and the run used about six billion output tokens.
The summary gives about 11 days. The method section of the same post says a little under two weeks. Kevin Buzzard, who leads the community effort to formalize Fermat's Last Theorem, endorsed the result.
OpenAI says it reached its automated research intern goal
OpenAI said on 6 September that it has met its September 2026 goal of an automated AI research intern, and it noted the agents still need human steering.
OpenAI defines an intern as a system that does well-defined research tasks, including ones that run for several days, under human direction. Its next target is an automated AI researcher by March 2028.
OpenAI gave two internal figures. The median OpenAI researcher was spending over $600 a day on inference at API prices by mid-August. Over the past six months, more than half of successful agent tasks lasting four to eight hours needed at least one human intervention.
The post does not say how often agents failed outright on those longer tasks.
Also in the news
- Gemini 3.8 Flash is Google's third Flash release in six weeks, with a company-reported 54.9% on HLE-Verified at $0.75 per million input tokens and $3.75 per million output tokens, and its Cyber variant goes to vetted defenders through the Fairwind Program.
- Muse Spark 1.3 from Meta uses about 20% fewer tool calls and 25% fewer tokens on coding than version 1.2, Meta reports, and Mark Zuckerberg again promised open weights "soon".
- WeatherNext 3 from Google DeepMind refreshes forecasts hourly from real-time satellite data at up to 5 km resolution and is built into Search, Gemini, Maps and Cloud.
- Atlas from World Labs is one model for camera-controlled video, 3D reconstruction and space simulation, and World Labs reports 75 to 94% human preference over five specialized video models on camera following.
- GWM Worlds 2 from Runway streams 720p video at 24 frames per second with synced audio, and several users can steer different characters in the same generated world.
- Solaris from Runway renders interactive software interfaces frame by frame in real time with no code behind them.
- Muse Voice Transcribe is Meta's first real-time audio perception model, doing streaming speech recognition and speaker separation in more than 70 languages.
- MAI-Transcribe-2 from Microsoft covers 60 languages with a 5.2% average word error rate on FLEURS, at a limited-time price of $0.10 per hour.
- MAI-Image-2.6 from Microsoft ranks second on Arena for text-to-image and editing, and Microsoft says its Flash variant is 2.8 times faster than GPT-Image-2-Medium.
- Qwen-Drive-1.0 from Alibaba is a 4B vision-language model for autonomous driving built on Qwen3.5-4B.
- Qwen3.8-Max-0902 is Alibaba's September snapshot of its hosted model, with better coding and multi-agent work.