The week in AI
The week in brief
Artificial Analysis scored Alibaba's new Qwen3-Max-Thinking at 40 on its Intelligence Index, behind DeepSeek V3.2 and GLM-4.7 at 42 and Moonshot AI's Kimi K2.5 at 47.
Chinese labs released most of the week's models. Alibaba shipped its closed reasoning flagship on 26 January. Moonshot AI released Kimi K2.5 with open weights on 27 January, and Meituan and StepFun followed with open mixture-of-experts models on 28 and 29 January.
Anthropic let MCP servers draw interactive interfaces inside Claude on 26 January. Google opened Genie 3 to paying subscribers as Project Genie on 29 January.
Alibaba's Qwen3-Max-Thinking scores 40 on Artificial Analysis
Alibaba released Qwen3-Max-Thinking on 26 January and claimed parity with GPT-5.2-Thinking, but Artificial Analysis placed it fourth among recent Chinese models.
Qwen3-Max-Thinking is a reasoning version of Qwen3-Max, Alibaba's model of more than a trillion parameters trained on 36 trillion tokens. Alibaba added large-scale reinforcement learning (RL) and a new test-time scaling scheme. The model decides for itself when to call tools while it reasons, and it doesn't wait to be told. It's a closed API with a 256,000-token context window.
Alibaba reports parity with OpenAI's GPT-5.2-Thinking. Artificial Analysis gave it 40 on its Intelligence Index, below DeepSeek V3.2 and GLM-4.7 at 42 and Kimi K2.5 at 47. The Humanity's Last Exam numbers don't line up. Alibaba's 58.3 used tools. Artificial Analysis measured 26% without tools.
Artificial Analysis lists the price as $1.20 per million input tokens and $6 per million output tokens for prompts up to 32,000 tokens. The price rises to $3 and $15 for prompts between 128,000 and 256,000 tokens.
Moonshot AI releases Kimi K2.5 with open weights
Moonshot AI released Kimi K2.5 on 27 January, an open-weight multimodal model with 1 trillion parameters that it says can coordinate up to 100 sub-agents in parallel.
Kimi K2.5 is a mixture-of-experts (MoE) model. Each token is routed to a small set of expert subnetworks, so only 32 billion of the 1 trillion parameters run at a time. Moonshot says it trained text and vision together from the start on about 15 trillion mixed tokens. It then fine-tuned with text-only data, which it calls zero-vision supervised fine-tuning.
The new feature is Agent Swarm. The model splits a task into pieces and runs them through many sub-agents at once. Moonshot trained this behaviour with RL on parallel agent runs, and it reports up to 4.5 times lower latency than a single agent on work that splits well. Moonshot also open-sourced a terminal coding tool, Kimi Code CLI.
Moonshot reports 76.8% on SWE-bench Verified, 96.1% on AIME 2025 and 78.4% on BrowseComp with Agent Swarm. The atlas has these figures from Moonshot's model card and has only partly confirmed the full release record. The independent number is the 47 that Artificial Analysis gave it, the highest of the Chinese models in its Qwen3-Max-Thinking comparison.
StepFun's open Step 3.5 Flash reports 74.4% on SWE-bench Verified
StepFun released Step 3.5 Flash on 29 January, an Apache-2.0 model with 196 billion total and 11 billion active parameters that it reports scores 74.4% on SWE-bench Verified.
StepFun built the model to make agent loops cheap. Three in every four attention layers look only at a sliding window of recent tokens, and the fourth looks at the full 256,000-token context. The model also predicts three tokens per step and doesn't stop at one, which StepFun calls MTP-3.
StepFun reports a typical speed of 100 to 300 tokens per second, with a peak of 350. Its paper lists 51.0% on Terminal-Bench 2.0, 88.2% on tau2-Bench, 86.4% on LiveCodeBench-v6 and 85.4% on IMO-AnswerBench. All of these are company figures. The RL setup covers math, code and tool use in one pipeline.
Meituan's LongCat team released a smaller open model the day before, on 28 January. LongCat-Flash-Lite has 68.5 billion total parameters and about 3 billion active, and more than 30 billion of those parameters sit in N-gram embedding tables. An N-gram embedding is a lookup table for short token sequences. Meituan reports that in some settings, making that table bigger beats adding more experts, and it puts less pressure on memory I/O. The model is MIT-licensed with 256,000 tokens of context through YaRN.
Anthropic lets MCP servers draw interactive apps inside Claude
Anthropic launched MCP Apps on 26 January, so any Model Context Protocol server can show an interactive interface inside a Claude conversation.
Until now, Model Context Protocol (MCP) servers gave Claude tools to call and returned text or data. MCP Apps let a server render a working interface in the chat. Users can edit a chart, move cards on a board, change a design or handle files without leaving the conversation.
The launch partners are Asana, Figma, Slack, Canva, Box and monday.com, and Salesforce is listed as coming soon. The feature works on web, desktop and mobile for every Claude plan, through the Claude directory.
On 30 January Anthropic added plugins to Cowork. A plugin bundles skills, connectors, slash commands and sub-agents, and Anthropic open-sourced 11 starter plugins for sales, finance, legal and biology.
Also in the news
- Project Genie opened Genie 3 to Google AI Ultra subscribers in the US aged 18 and over on 29 January as an experimental prototype for creating, exploring and remixing interactive worlds, with a 60-second cap that Google puts down to compute limits and without the promptable events shown in the August 2025 demo.
- Lucy 2.0 from Decart, released on 26 January, edits live video at 30 frames per second and 1080p with a pure diffusion model and no depth maps or 3D, for swapping characters, clothing, products and backgrounds.
- DeepSeek-OCR 2 is a 3-billion-parameter document reader released on 27 January whose encoder orders image patches by meaning instead of scanning them row by row.
- Qwen3-ASR arrived from Alibaba on 29 January as open speech recognition in 0.6-billion and 1.7-billion sizes for 52 languages and dialects, with a forced aligner that timestamps up to five minutes of speech.
- Ray3.14 from Luma Labs, released on 26 January, generates native 1080p video, and Luma says it's four times faster and three times cheaper than Ray3 at 720p.
- Prism is a free OpenAI workspace for scientists to write and collaborate on research, running on GPT-5.2, launched on 27 January.
- Mistral Vibe 2.0 made Mistral's terminal coding agent a paid product on 27 January, with custom subagents, slash-command skills and multiple-choice clarifying questions.
- Grok Imagine API opened on 28 January, giving developers xAI's text-to-video, image-to-video and video editing with native audio.
People
- Toby Pohlen, an xAI co-founder, left xAI. The record dates his departure to February 2026 and gives no exact day.