The week in AI

Week of 29 Jun to 5 Jul 2026

The week in brief

Meituan released LongCat-2.0 on 29 June, a 1.6 trillion parameter open-weight model trained entirely on Chinese accelerators, which scored 59.5 on SWE-bench Pro by its own measure.

Meituan's LongCat team published weights for its new model under the MIT licence on 5 July. Anthropic shipped Claude Sonnet 5 and a science workbench called Claude Science on 30 June. On 2 July it published details of the cyber safeguards on Fable 5.

Mistral released Leanstral 1.5 on 2 July, an open-weight theorem prover that it reports solves every problem in miniF2F. OpenAI lost two senior safety leaders in July, Joshua Achiam and Johannes Heidecke.

Meituan released LongCat-2.0, trained on 50,000 Chinese chips

Meituan's LongCat-2.0 is a 1.6 trillion parameter mixture-of-experts model with a 1 million token context, released under MIT and trained on more than 35 trillion tokens using only Chinese AI accelerators.

A mixture-of-experts (MoE) model holds many expert subnetworks and sends each token through only a few of them, so the total parameter count is far larger than the compute spent per token. Meituan says pretraining ran on more than 35 trillion tokens with no rollbacks. According to VentureBeat, it used more than 50,000 Chinese ASICs, and Meituan says it used a Huawei communication library.

The architecture has two features worth knowing. LongCat Sparse Attention builds on the indexer from DeepSeek's DSA, a small component that picks which earlier tokens each token should attend to so long contexts cost less. The model also carries 135 billion parameters of N-gram embeddings, which sit alongside the experts.

Meituan reports 59.5 on SWE-bench Pro, against its own figures of 58.6 for GPT-5.5 and 69.2 for Claude Opus 4.8. It also reports 70.8 on Terminal-Bench 2.1 and 79.9 on BrowseComp. These are company numbers, and no independent evaluation appeared that week.

Before launch the model ran anonymously on OpenRouter as "Owl Alpha", where it handled about 559 billion tokens a day. VentureBeat reports API pricing of $0.30 per million input tokens and $1.20 per million output tokens during a promotion, and $0.75 per million input tokens and $2.95 per million output tokens at standard rates. Weights went up on Hugging Face on 5 July.

Anthropic released Claude Sonnet 5 and Claude Science on 30 June

Anthropic released Claude Sonnet 5 on 30 June as the default model for Free and Pro users, at an introductory price of $2 per million input tokens and $10 per million output tokens.

Anthropic says Sonnet 5 is the most agentic Sonnet so far, with large gains over Sonnet 4.6 in reasoning, tool use, coding and knowledge work. At higher effort settings it matches Opus 4.8 on some tasks, by Anthropic's account. It has a 1 million token context and uses adaptive thinking by default. The list price had been $3 per million input tokens and $15 per million output tokens, and the $2 and $10 introductory price was made permanent on 10 August.

On safety, Anthropic reports a lower rate of undesirable behaviours than Sonnet 4.6 and far weaker cyber ability than the Opus models. That second point matters for anyone comparing it to Fable 5, where Anthropic has put cyber classifiers in front of the model.

The same day Anthropic opened a beta of Claude Science, a workbench for research analysis. A coordinating agent with more than 60 skills and connectors runs the work, and a separate reviewer agent checks it. Every figure comes with the exact code, environment and message history that produced it, so an analysis can be rerun. It runs locally on macOS or Linux, over SSH, or on a cluster node, renders 3D proteins and genome tracks, and can move work from a laptop to GPUs on demand. It's available on Pro, Max, Team and Enterprise plans.

On 2 July, after redeploying Fable 5, Anthropic published a list of the harm types its cyber classifiers target and the ones they don't block. It also proposed a draft scale for grading how severe a jailbreak is, built with Amazon, Microsoft, Google and other Glasswing partners, so labs and governments can describe a bypass in shared terms. Anthropic presented the scale as an early draft for discussion.

Mistral released Leanstral 1.5, which it says solves all of miniF2F

Mistral reports that Leanstral 1.5, a 119 billion parameter model with 6 billion active, scores 100% on miniF2F and solves 587 of 672 PutnamBench problems, with weights under Apache 2.0.

Leanstral writes proofs in Lean, a language where a proof is a program that a checker either accepts or rejects. That makes the scores pass or fail with no grading judgement involved. Mistral reports 100% on both the validation and test splits of miniF2F, a set of competition maths problems, which means that benchmark no longer separates models.

PutnamBench is harder and draws on the Putnam undergraduate competition. Mistral says Leanstral 1.5 solved 587 of its 672 problems with a 4 million token budget. It also reports 87% on FATE-H and 34% on FATE-X, and says the model found five previously unknown bugs across 57 open-source repositories. All of these figures come from Mistral.

Also in the news

  • Cognition launched Devin Fusion on 29 June, which splits work between a frontier lead model and a cheaper sidekick model, and claims Fable 5-level quality at 60% lower cost.
  • xAI reportedly released Grok Imagine video 1.5 in July, with text, image and reference-to-video modes at native 1080p, priced at $0.08 per second in the API.
  • Mistral AI CEO Arthur Mensch reportedly wrote that a "fat but sparse" open-weight MoE family, larger in total parameters than the 675 billion parameter Large 3, would enter partner early access in July, according to TechTimes. Mistral disclosed no parameter count, benchmarks, licence or date.

People

  • Joshua Achiam left OpenAI in July after nearly nine years. OpenAI had disbanded his mission alignment team earlier in 2026 and made him chief futurist.
  • Johannes Heidecke, head of Safety Systems at OpenAI since 2024, left in July in a reorganization that put OpenAI's safety teams under Mia Glaese, VP of research and safety.