The week in AI
The week in brief
MiniMax released MiniMax M3 on 1 June, an open-weights model that it reports scores 80.5% on SWE-bench Verified and serves 1 million token contexts at about a twentieth of M2's per-token cost.
MiniMax led the week with M3 and a sparse attention design that cuts the cost of very long contexts. Microsoft announced MAI-Thinking-1, its first in-house reasoning model, at Build on 2 June. Google released Gemma 4 12B on 3 June.
Two coding tool companies shipped desktop apps built for running many agents at once. Cognition launched Devin Desktop on 2 June, and GitHub followed with a standalone Copilot app on 3 June.
MiniMax releases M3 with sparse attention for 1M context
MiniMax M3 is a natively multimodal open-weights model with about 428 billion total parameters, and MiniMax reports 80.5% on SWE-bench Verified.
M3 is a mixture-of-experts model, so only about 23 billion of its parameters are active for each token. It handles a context of 1 million tokens. MiniMax says each token costs about a twentieth of what it cost on M2, the previous generation.
The main change is MiniMax Sparse Attention (MSA). In full attention, each new token looks at every earlier token. MSA instead picks the blocks of stored keys and values that matter and attends only to those. MiniMax reports a 9x speedup in prefill and a 15x speedup in decoding against M2 at 1 million tokens. A separate arXiv paper credits MSA with cutting attention compute by 28.4x.
MiniMax reports 59.0% on SWE-bench Pro and 66.0% on Terminal-Bench 2.1. On the multimodal side it reports 78.1% on MMMU Pro. All of these figures come from MiniMax, and the records for this week contain no independent evaluation.
The weights are on Hugging Face under a community licence that places restrictions on use. That makes M3 available to self-host, with limits that depend on the licence terms.
Microsoft announces MAI-Thinking-1, its first in-house reasoning model
At Build on 2 June, Microsoft announced MAI-Thinking-1, a sparse mixture-of-experts reasoning model with about 1 trillion total parameters and 35 billion active.
MAI-Thinking-1 has a 256,000 token context. Microsoft says it trained the model without third-party distillation, which means it didn't learn from another company's model outputs. The model is available through a closed API.
Microsoft reports 97.0% on AIME 2025 and 94.5% on AIME 2026. It says the model matches Claude Opus 4.6 on SWE-Bench Pro. In Microsoft's blind human evaluations across 1,276 single-turn and multi-turn tasks, raters preferred MAI-Thinking-1 to Claude Sonnet 4.6. These are all Microsoft's own numbers.
Microsoft announced MAI-Thinking-1 alongside other MAI models. MAI-Code-1-Flash is an agentic coding model with 5 billion active parameters. It ships in GitHub Copilot and VS Code, and Microsoft describes it as Haiku-class at a lower cost.
GitHub and Cognition ship desktop apps for running many agents
GitHub released a standalone Copilot desktop app in technical preview on 3 June, a day after Cognition launched Devin Desktop.
The GitHub Copilot app is built for running many agents in parallel. Each agent works in its own git worktree, a separate working copy of the repository, so agents don't overwrite each other's changes. A "My Work" view gathers sessions, issues, pull requests and automations in one place.
The app also has Canvas surfaces, local and cloud sandboxes, and remote control from a phone. It requires a Copilot Pro, Pro+, Business or Enterprise plan. It follows the same design as Cursor 3 and the Codex app, with the agent as the main surface and the editor secondary.
Cognition's Devin Desktop grew out of Windsurf, the editor Cognition already ships, and Cognition describes it as an evolution of Windsurf. Agent Command Center is the default view, and the app adds Spaces. It stays compatible with Windsurf.
Devin Desktop supports the Agent Client Protocol (ACP), a standard way for an editor to talk to coding agents. ACP lets agents from other companies run inside Devin Desktop next to Devin.
Google releases Gemma 4 12B with no separate vision or audio encoders
Google DeepMind released Gemma 4 12B on 3 June, an open 12 billion parameter model under Apache 2.0 that runs in 16GB of memory.
Most multimodal models pass images and audio through separate encoder networks before the language model sees them. Gemma 4 12B sends vision and audio inputs straight into the language model, a design Google calls "unified, encoder-free". It is the first mid-sized Gemma with native audio input.
Google says the model performs close to the 26B mixture-of-experts Gemma 4 while using less than half the total memory. It ships with multi-token prediction (MTP) drafters, which guess several tokens ahead so the model can check them in one step and generate faster.
Google also reported that Gemma 4 models had passed 150 million downloads.
Also in the news
- Project Glasswing expanded on 2 June to about 150 more organizations in over 15 countries, adding power, water, healthcare and communications operators plus hardware vendors. Anthropic said it expects many other AI companies to have Mythos-class models within 6 to 12 months, possibly without safeguards.
- MAI-Image-2.5 is Microsoft's new text-to-image and editing model, with a Flash version. Microsoft says it beats Google's Nano Banana Pro on Arena Elo, and it is in PowerPoint and Foundry.