The week in AI
The week in brief
Moonshot AI released Kimi K3 on 16 July, a mixture-of-experts model with 2.8 trillion parameters that was reported third on the Artificial Analysis leaderboard, behind two closed models.
Kimi K3 was the week's main release, and Moonshot launched it through its API first. Alibaba's Qwen team started a new speech line on 14 July with Qwen-Audio-3.0 text-to-speech. OpenAI described GPT-Red, a red-teaming model it uses to harden GPT-5.6, and Decart updated its live video editor.
Moonshot AI launched Kimi K3, a 2.8-trillion-parameter model
On 16 July Moonshot AI released Kimi K3, which has 2.8 trillion total parameters and 104 billion active per token, with a 1 million token context window.
Kimi K3 is a mixture-of-experts (MoE) model. Each layer holds many small expert networks, and a router sends every token to only a few of them, so the compute per token is a fraction of the total size. K3 has 93 layers with 896 experts each, and 16 are active per token. Moonshot reports about 2.5 times better scaling efficiency than Kimi K2.
The attention design is new for Moonshot. Kimi Delta Attention (KDA) is a hybrid that mixes cheaper attention layers with standard ones to keep very long contexts affordable, and the model also uses a component Moonshot calls Attention Residuals. The weights are stored in MXFP4, a 4-bit floating point format. Moonshot trained with quantization in the loop from supervised fine-tuning onward, so the model learned to work at that low precision instead of being compressed after training.
Moonshot reports 93.5 on GPQA-Diamond. On Terminal-Bench 2.1 it reports 88.3, run in its own Kimi Code harness. Its BrowseComp score is 91.2 when the context is compacted at 300,000 tokens, and 90.4 when the model uses the full 1 million tokens with no context management. All of these are company-reported.
Moonshot doesn't claim the top spot. It says K3 still trails Claude Fable 5 and GPT-5.6 Sol, which fits the reported third place on Artificial Analysis. The API costs $3 per million input tokens and $15 per million output tokens, with cached input at $0.30 per million tokens. At launch the model was available only through that API.
Also in the news
- Qwen-Audio-3.0 arrived on Alibaba Cloud on 14 July as a text-to-speech model in Plus and Flash versions, and Alibaba says the Flash version sends its first audio packet in 200 milliseconds.
- GPT-Red is an OpenAI model trained by self-play to attack OpenAI's own models, and OpenAI said on 15 July that it uses those attacks to train GPT-5.6 against prompt injection.
- Lucy 2.5 from Decart, released on 16 July, edits live video at 30 frames per second, adding physically aware effects such as water, sand and fire, object removal and whole-scene style changes, with better consistency from frame to frame.