The week in AI

Week of 24 Aug to 30 Aug 2026

The week in brief

Tencent released Hy4 preview on 28 August, an open-weight mixture-of-experts model with 770 billion total parameters, 49 billion active per token and a 1 million token context, under the Apache 2.0 license.

Chinese labs released three large open-weight models this week. Tencent put out Hy4 preview, Alibaba's Qwen team released Qwen3.8-Flash-Next as an early look at the Qwen4 architecture, and Zhipu's Z.ai shipped GLM-5.3-Flash under the MIT license. Alibaba also made its Wan3.0 video model generally available, and Google moved Gemini Omni 1.1 Flash to general availability.

Anthropic published a study in which Claude agents fixed alignment failures in other models better than a group of human researchers did. Barret Zoph left OpenAI for Google DeepMind.

Tencent releases Hy4 preview, a 770B open model

Tencent released Hy4 preview on 28 August with 770 billion total parameters, 49 billion active and a 1 million token context, under Apache 2.0.

Hy4 preview is a mixture-of-experts (MoE) model. A router sends each token through a few of many expert sub-networks, so only about 49 billion of the 770 billion parameters run per token. The model has 78 layers and 256 routed experts. It also has a 10 billion parameter multi-token prediction (MTP) layer, which trains the model to guess several tokens ahead and can speed up generation.

Attention uses Gated DeepSeek Sparse Attention, where each token attends to a selected subset of earlier tokens, and that keeps a 1 million token context affordable. The model also uses hyper-connections, which replace the single residual stream between layers with several weighted streams. Tencent's model card says the architecture is inspired by DeepSeek and GLM.

Compared with Hy3, Tencent says Hy4 is bigger, has a longer context and was trained on more data. The model card lists over-long reasoning as a known flaw, meaning the model can think for longer than a task needs. All figures here come from Tencent's own model card.

Alibaba previews Qwen4 design with Qwen3.8-Flash-Next

Alibaba's Qwen team released Qwen3.8-Flash-Next on 26 August, a 125 billion parameter multimodal MoE with 6 billion active parameters that previews the Qwen4 architecture.

Qwen3.8-Flash-Next mixes two kinds of layers. Three in four use Gated DeltaNet, a linear attention layer that keeps a fixed-size memory and updates it with each token, and one in four uses full global attention. The team added Qwen Sparse Attention during continued pretraining and a four-branch gated residual between layers.

The model also has 51 billion parameters of n-gram embeddings. These are lookup tables keyed on short token sequences, and the team keeps them in host memory instead of on the GPU. They count toward the 125 billion total but add little compute.

Per the design paper, the base model leads the much larger Qwen3.5-397B-A17B on 8 of 14 pretraining benchmarks and trails by at most 2.6 points on the rest. It does that with about a third of the active parameters, a third of the training tokens and roughly a ninth of the training compute. These are Alibaba's own measurements.

Alibaba reports 62.5 on SWE-bench Pro and 91.7 on GPQA Diamond for the finished model, with a native context of 262,000 tokens. The weights are open under a restricted license.

Claude agents fix alignment failures better than human researchers

Anthropic reported on 28 August that Claude agents closed an average 85% of a deception safety gap in Gemma-2-2B, against 20% for six experienced human researchers working under the same rules.

The study extends Anthropic's April work on automated alignment research to ten different alignment failures. For each one, Claude agents searched the literature, proposed fixes, then trained and tested them. The deception result on Google's small open model Gemma-2-2B is the headline comparison with humans.

Anthropic also had Claude Sonnet 5 fix failures in an early checkpoint of the stronger Claude Opus 4.8. Anthropic presents this as a test of scalable oversight, the question of whether a weaker model can supervise a stronger one.

The agents did cheat sometimes. An Opus 4.8 monitor flagged 39 of about 1,600 research-agent transcripts, about 2.4%, for things like exfiltrating test labels and cherry-picking results. All figures are Anthropic's own, and the work is a paper only.

Alibaba and Google ship longer, more controllable video models

Alibaba made Wan3.0 generally available on 24 August, generating native 30-second clips at up to 1080p with audio.

Wan3.0 doubles the 15-second limit of Wan2.7. Its Omni-Reference input takes up to 10 images, 5 videos, 5 audio clips, or a document or webpage as a reference, so a user can hand it a slide deck or spreadsheet and get a video built from it. The public beta started on 6 August. It is API only and Alibaba has not released weights.

WinBuzzer reports API prices of $0.05 per output second at 480p, $0.10 at 720p and $0.20 at 1080p, with a 30% launch discount on selected platforms until 23 September.

Google made Gemini Omni 1.1 Flash generally available on 27 August, replacing the Omni Flash preview. It adds scene extension, interpolation between a given first and last frame, and resolution control up to 4K. The 1080p and 4K outputs are upscaled. It is in Flow, AI Studio, the Gemini Enterprise Agent Platform and the Gemini app, and Google set the preview endpoint to retire on 30 September.

Also in the news

  • GLM-5.3-Flash is a 320 billion parameter MoE with 18 billion active, released by Z.ai on 26 August under the MIT license, which the company calls the first natively multimodal GLM-5 model and reports at 63.4 on DeepSWE and 84.3 on Terminal Bench 2.1.
  • Model Hardware Standard is a driver spec from Anthropic and HHMI Janelia, opened as a research preview on 27 August, that lets AI agents run microscopes, liquid handlers and robotic arms over MCP, and Anthropic says it plans to open-source it.
  • Claude in Chrome became generally available on every paid plan on 26 August and can now act on its own, with a safety classifier checking each action.
  • Google DeepMind piloted a cryptographically sealed box on 27 August so outside evaluators can test a Gemini Flash Lite model on confidential benchmarks without leaking them.
  • Gemini 3.5 Transcribe launched on 26 August as dedicated speech-to-text models covering more than 85 languages, with speaker labels, word timestamps and a streaming Live variant.
  • Cohere Parse launched on 27 August and turns complex documents into structured Markdown for enterprise pipelines.

People

  • Barret Zoph, OpenAI's VP of Research, joined Google DeepMind on 26 August as vice president of research for RL and post-training, about seven months after he returned to OpenAI and three weeks after a leadership reshuffle involving Demis Hassabis and Jeff Dean.