The week in AI
The week in brief
SpaceX closed its $60 billion all-stock acquisition of Anysphere, the company behind the Cursor coding editor, on 14 August.
Cursor is now a wholly owned subsidiary inside SpaceXAI, and two days earlier xAI's Grok 4.6 went live in Cursor on all plans. On 10 August Anthropic published a number theory result produced by an unreleased Claude model, which raised the proven share of Riemann zeta zeros on the critical line from 41.6% to 67.2%.
The rest of the week was open weights. Meta released Muse Glimmer 30B under Apache 2.0, its first open-weights model since Llama 4. Alibaba, DeepSeek and Zhipu each shipped large open coding models, and Google released Gemini 3.7 Flash.
SpaceX closes $60 billion Cursor deal, ships Grok 4.6 in it
SpaceX completed its purchase of Anysphere on 14 August, paying in stock and folding Cursor into its SpaceXAI unit.
Cursor's shares converted into about 389.3 million SpaceX Class A shares, or about 391 million once restricted stock units and options are counted. Cursor keeps its name for now. Reports say the Cursor brand is expected to give way to Grok branding eventually, but no date has been given.
The first product of the combined company came before the close. On 12 August xAI released Grok 4.6, a model tuned for long-running agents that do multi-step research, analysis and app-building. It is available in Cursor on every plan, and through Grok Build, the xAI API, OpenRouter, Vercel and Cloudflare.
xAI reports a score of 61 on the Artificial Analysis Intelligence Index, up from 56 for Grok 4.5 and level with GPT-5.6 Sol by xAI's count. xAI also reports 65.9% on DeepSWE v1.1 and 69.9% on CursorBench v3.2, Cursor's own coding benchmark. All of these figures come from xAI's release table.
The price didn't change. For prompts under 200,000 tokens it is $2 per million input tokens and $6 per million output tokens.
Unreleased Claude raises Riemann zeta bound to 67.2%
Anthropic says an unreleased Claude model proved that at least 67.2% of the Riemann zeta function's nontrivial zeros lie on the critical line, up from the previous bound of 41.6%.
The Riemann hypothesis states that every nontrivial zero of the zeta function lies on one vertical line, the critical line. Nobody has proved it. Mathematicians have instead proved lower bounds on how many of the zeros must sit there, and the best published bound before this was 41.6%.
Claude was set to attempt the hypothesis itself and failed. On the way it combined three existing pieces of work into the new bound. One is by Aryan, one is by Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh, and one is a 2000 paper by Enrico Bombieri.
Two mathematicians at Anthropic checked the argument and wrote an expert note on it. Claude also produced a proof in a form that can be formally verified by a proof checker. Brian Conrey and Dan Goldston reviewed the result on short notice before publication.
The model is a research model and Anthropic hasn't released it. Access is through the paper only.
Meta releases Muse Glimmer 30B under Apache 2.0
On 10 August Meta published the weights of Muse Glimmer, a 30 billion parameter dense model under the Apache 2.0 licence and its first open-weights release since Llama 4.
Meta had kept its recent Muse models closed, and Glimmer partly reverses that. The licence is permissive, the weights are on Hugging Face, and trade press reports that the model runs on a single 24 GB GPU when quantised.
Glimmer was trained by logit distillation from Muse Spark, Meta's larger closed model. The small model learns to match the full probability Spark assigns to each possible next token, and Meta then added reinforcement learning (RL) on top.
The model is built for agents running locally. It includes a perception encoder for image input and a speculative-decoding drafter, a small companion model that guesses several tokens ahead so the main model can check them in one pass. Meta compares it with Gemma4-31B and Qwen3.6-27B on agentic, coding and multimodal suites.
Alibaba, DeepSeek and Zhipu ship open coding models
Three Chinese labs released open-weights coding models between 12 and 14 August, led by Alibaba's 2.4 trillion parameter Qwen3.8.
On 12 August Alibaba's Qwen team released Qwen3.8-2.4T-A95B. It is a mixture-of-experts model, so each token is routed to 10 of 512 expert blocks plus one shared expert, and about 95 billion parameters are active per token. It mixes Gated DeltaNet layers, a cheaper linear form of attention, with gated full attention across 92 layers. Thinking mode is always on, with three effort settings.
Alibaba reports 86.6 on Terminal Bench 2.1 and 67.7 on SWE-bench Pro. The weights come under a Qwen3.8-Max licence and not Apache 2.0. Two days later Alibaba released Qwen3.8-27B under Apache 2.0, a dense model with image and hour-scale video understanding that Alibaba reports at 61.7% on SWE-bench Pro and 84.3% on OSWorld-Verified.
On 13 August DeepSeek made DeepSeek-V4-Pro-0813 generally available as an open-weights model, with post-training aimed at agents. DeepSeek reports 87.9 on Terminal Bench 2.1 and 62.7 on DeepSWE. The API now speaks OpenAI's Responses format natively so it works with Codex, and from 16 August it charges half price off-peak. Output costs $3.96 per million tokens at peak and $1.98 off-peak.
On 14 August Zhipu released GLM-5.3, a coding and security post-train of the GLM-5.2 base. Zhipu reports DeepSWE v1.1 rising from 46.2 to 66.9 and 84.5% on CyberGym, which Zhipu calls the top score on that benchmark. The Decoder reports Zhipu's claim that the model found 2,436 vulnerabilities across 269 projects. Zhipu says it ran its most extensive risk review before releasing the weights, and it replaced the MIT licence of earlier GLM models with a custom GLM-5.3 licence.
Also in the news
- Gemini 3.7 Flash came out on 13 August, three weeks after 3.6 Flash, and Google reports 65.3% on DeepSWE v1.1, up from 49.0%, at an introductory $0.75 per million input tokens and $3.75 per million output tokens through 31 December 2026.
- Lovable raised $400 million at a $13.3 billion valuation on 12 August, led by Menlo Ventures and EQT's Scaleup Europe Fund, and says more than 60 million projects have been built on it since its November 2024 launch.
- DeepSeek Harness is an MIT-licensed, plugin-based agent framework that DeepSeek released as a developer preview on 13 August, with web, desktop and SSH interfaces, subagents and MCP support.
- Suno signed a global licensing deal with BMG on 12 August that covers recordings and publishing and settles past use, ahead of its first model built with industry partners.
- Wan 3.0 from Alibaba's Tongyi lab generates up to 30 seconds of video in one pass from text, images, audio, video or documents, priced from $0.05 per second at 480p to $0.20 per second at 1080p.
- MiniMax Music 3.0 is an open-weights model that composes full songs of up to five minutes, built from an 8 billion parameter language model initialised from Qwen3.5-8B plus a flow-matching renderer, with no benchmarks published.
- Anthropic said on 14 August that future Claude models will carry a statistical text watermark, adopted with other providers to meet the EU AI Act's marking rule.
- MAI-Code-1.1-Flash from Microsoft costs a quarter as much as the previous version, and Microsoft reports it is 22% better on Terminal-Bench 2.1 in Copilot CLI.
- MAI-Cyber-1-Flash, Microsoft's security model inside its MDASH multi-agent system, scores 96% any-crash on CyberGym at half the earlier cost, by Microsoft's count.
- Higgsfield released Cinema Studio 4.0 with 30-second generations and up to 50 references per shot.
People
- Brad Lightcap said on 11 August that he is leaving OpenAI after eight years to start something new, after moving from chief operating officer to special projects in April.
- Lin Junyang, the former Qwen technical lead, launched Shanghai-based Pragmatik Labs on 12 August to build agents for digital and physical tasks, with an angel round co-led by Gaorong and HSG and backing from Tencent.