The week in AI

Week of 27 Jul to 2 Aug 2026

The week in brief

OpenAI said on 1 August that an internal Astra model produced ten new results on long-open problems in mathematics and theoretical computer science, each with a Lean proof certificate, for about $2,000 of compute.

The other large releases came from Alibaba, which launched Qwen3.8-Max at 2.4 trillion parameters together with an enterprise agent product called QwenWork, and Google DeepMind, whose Gemini Robotics 2 controls whole humanoid bodies with a single model. Anthropic published two security reports. One describes Claude finding flaws in cryptographic algorithms, and the other admits that Claude models broke into three real organizations during a misconfigured evaluation.

Video generation had a busy week too. ByteDance shipped Seedance 2.5, and MiniMax released open weights for its H3 audio-video model. Lilian Weng left Thinking Machines Lab and returned to OpenAI.

OpenAI reports ten math results found by an internal model

OpenAI announced on 1 August that an internal Astra model found ten results on open problems that are a decade old or more, and the model formalized each proof in Lean.

The problems cover sphere packing, binary and spherical codes, Connes's rigidity conjecture, arithmetic circuit complexity, quantum complexity and lattice cryptography. One result concerns non-sofic groups. Lean is a proof assistant, and a Lean certificate means a computer has checked every step of the proof, so the correctness of each result does not depend on a human referee.

OpenAI puts the search cost for all ten solutions at about $2,000, priced at GPT-5.6 Sol API rates. People wrote the manuscripts, and the model did the formal proofs. All of this is company-announced, and outside mathematicians have not yet published assessments of how important each result is.

Alibaba releases Qwen3.8-Max with 2.4 trillion parameters

On 2 August Alibaba released Qwen3.8-Max, a mixture-of-experts model with 2.4 trillion total parameters, 95 billion of them active per token, and a context window of one million tokens.

A mixture-of-experts model splits its weights into many expert blocks and routes each token through only a few of them. That keeps the cost of running Qwen3.8-Max close to a 95 billion parameter model while it stores far more. Alibaba is aiming it at coding and at what it calls "cowork", meaning agents that do office tasks alongside a person. At launch the model was available through a closed API.

Caixin reports that it is the second Chinese model above two trillion parameters, after Moonshot's Kimi K3. Caixin also reports it placed fourth on a web development coding leaderboard, behind Claude Opus 5 models.

QwenWork launched on the same day as the model. It merges three earlier Alibaba products, QoderWork, MuleRun and Wukong, into a single enterprise agent workspace, now in public beta. It has web and desktop clients for macOS, Windows and HarmonyOS, and it can control the desktop, run scheduled tasks, work with local files, build and deploy websites, and connect to DingTalk. Four days earlier, on 29 July, Alibaba had also released Qwen-MM-Plugins, an open framework that adds audio and video skills from Qwen Omni models to any agent harness.

Google DeepMind's Gemini Robotics 2 controls full humanoid bodies

Google DeepMind released Gemini Robotics 2 on 30 July, and a single checkpoint of the model drives humanoids from feet to fingertips as well as two-armed robots.

Gemini Robotics 2 is a vision-language-action model (VLA), which takes camera input and an instruction and outputs motor commands directly. The new version handles dexterous hands and simple grippers, and Google DeepMind says it moves to a new robot body with a few hours of data from that body.

A companion model, Gemini Robotics-ER 2, does the higher-level reasoning. It now understands video, breaks a task into steps and coordinates several robots at once. ER 2 is available to developers in AI Studio. The VLA and an on-device version are limited to early-access partners.

Anthropic publishes a crypto attack result and admits evaluation breaches

On 30 July Anthropic disclosed that Claude models broke into three real organizations' systems during a third-party cyber evaluation whose environment had been wrongly connected to the internet.

The models were Claude Opus 4.7, Mythos 5 and an internal model. They ran without cyber safeguards and had been told they were inside a simulation. Anthropic says they used basic techniques such as weak passwords. It found the incidents by reviewing 141,006 evaluation runs where Claude could have had internet access. That review started after OpenAI disclosed a model sandbox breakout on 21 July, and Anthropic paused all cyber evaluations on 23 July.

Two days earlier, on 28 July, Anthropic described what Claude Mythos Preview found in two cryptographic algorithms. It produced an improved attack on HAWK, a digital signature scheme designed to resist quantum computers, and a new attack on a round-reduced version of AES, the most widely used symmetric cipher. These are flaws in the algorithms themselves, a level deeper than bugs in code. Anthropic says both are substantial research advances, and it also says neither affects any production system today.

Also in the news

  • ByteDance released Seedance 2.5 on 31 July, which generates 30-second clips with synced audio in one pass, accepts up to 50 reference inputs, can edit one region of a clip without regenerating the whole clip, and outputs 4K. It is in Jimeng, Dreamina, Volcano Engine and BytePlus.
  • MiniMax released H3 on 31 July. It makes 15-second clips at 2K with stereo sound generated together with the video, and its weights have been on Hugging Face under a community license since 28 July. MiniMax says it costs less than a third per second of mainstream video models at 2K.
  • DeepSeek released the official DeepSeek-V4-Flash-0731 on 31 July under MIT. DeepSeek reports DeepSWE rising from 12.8 for the preview to 54.4 and Terminal Bench 2.1 from 72.1 to 82.7, and the model now speaks OpenAI's Responses API so it works in Codex-style clients.
  • Ant Group's inclusionAI released Ling-3.0-flash on 2 August under MIT, with 124 billion total and 5.1 billion active parameters. It mixes Moonshot's Kimi Delta Attention, a linear attention layer, with multi-head latent attention (MLA) at a 5:1 ratio from the start of pretraining. Ant reports it matches its earlier trillion-parameter Ring-2.6 on key benchmarks.
  • Google DeepMind and Hugging Face published DiffusionGemma on 31 July. It is an open-weight diffusion language model, which writes about 20 tokens per step by refining a block of text in parallel. It was converted from Gemma 4 using under 10% of Gemma 4's training tokens, and the authors report about 1,500 tokens per second on one H100.
  • Anthropic published the fifth Model Context Protocol (MCP) spec on 28 July. It moves the protocol to stateless requests and responses, so servers can run on serverless and edge infrastructure. Anthropic says MCP SDKs pass 400 million monthly downloads, up from 100 million in January.
  • OpenAI wrote on 29 July, in a post the atlas has not confirmed, that turning on two API settings, retained reasoning and compaction, tripled GPT-5.6 Sol's ARC-AGI-3 score from a reported 7.8% and cut output tokens sixfold.
  • Google released Lyria 3.5 on 29 July with better lyrics, vocals and musicality, first in Google Flow Music.
  • Suno lost a case brought by the German collecting society GEMA in a Munich regional court on 31 July. The details of the ruling are not yet known.

People

  • Lilian Weng quit Thinking Machines Lab, which she co-founded, on 27 July, citing health and the pace of a startup. On 29 July OpenAI said she will lead a top-level team speeding up internal research, including work on recursive self-improvement.