The week in AI

Week of 6 Jul to 12 Jul 2026

The week in brief

OpenAI released GPT-5.6 on 9 July in three tiers, and it reports that the top tier, Sol, scores 53.6 on Agents' Last Exam.

OpenAI also launched ChatGPT Work, an agent for office tasks, and replaced ChatGPT's voice mode with a new model called GPT-Live. On 6 July Tencent released the finished Hy3 under an open licence, with API input priced at 1 yuan per million tokens. Meta shipped its first in-house image model and opened a paid API for its Muse models, and xAI released Grok 4.5, a coding model trained with Cursor.

OpenAI splits GPT-5.6 into Sol, Terra and Luna

On 9 July OpenAI released GPT-5.6 as three models at three prices, after a limited preview that began on 26 June.

Sol is the top tier. OpenAI reports it scores 53.6 on Agents' Last Exam, which OpenAI says is 13.1 points above Claude Fable 5, and 92.2% on BrowseComp. OpenAI also says Sol reaches these scores with fewer tokens than earlier models. All of these are OpenAI's own results.

Launch prices are $5 per million input tokens and $30 per million output tokens for Sol. Terra costs $2.50 per million input tokens and $15 per million output tokens, and Luna costs $1 per million input tokens and $6 per million output tokens. The API adds Programmatic Tool Calling, explicit prompt-cache breakpoints and a max reasoning setting. It also adds an "ultra" mode that runs four agents in parallel, and a Multi-agent beta.

The safety card has a large jump in offensive cyber ability. OpenAI reports Sol scores 73.5% on ExploitBench with production safeguards switched off, against 47.9% for GPT-5.5. OpenAI says its cyber safeguards now block about ten times more harmful activity.

ChatGPT Work launched alongside GPT-5.6. It's an agent built on Codex technology that works across apps and files for hours and returns finished spreadsheets, slides, documents and web apps, with Scheduled Tasks and Sites. OpenAI says more than 5 million people use Codex each week, and more than 1 million of them use it outside software development. ChatGPT Work is on paid plans, and free and Go users get Luna in the desktop app.

Tencent releases the finished Hy3 under Apache 2.0

On 6 July Tencent released the completed Hy3 as open weights under the Apache 2.0 licence, with API input at 1 yuan per million tokens.

Hy3 keeps the architecture of the earlier version. It's a mixture-of-experts model with 295 billion parameters in total and 21 billion active for each token. In a mixture-of-experts model, a router sends each token to a few specialised sub-networks, so only part of the model runs at a time. The API costs 1 yuan per million input tokens and 4 yuan per million output tokens.

Tencent fed post-training with feedback from more than 50 of its own products, including Yuanbao and WorkBuddy, where Hy3 is already built in. Tencent reports that hallucinations on its internal evaluations fell from 12.5% to 5.4%. In a blind test that Tencent ran with 270 experts scoring on a scale of 4, Hy3 scored 2.67 and GLM-5.1 scored 2.51.

On SWE-Bench Pro, Tencent reports 57.9 for Hy3 and 69.2 for Claude Opus 4.8. That gap to the closed frontier on coding is in Tencent's own figures.

Meta ships Muse Image and opens a paid Muse API

Meta released Muse Image, its first in-house image model, on 7 July, and opened a public preview of the Meta Model API with Muse Spark 1.1 on 9 July.

Muse Image works as an agent. It can write code, search the web and critique its own output before editing it, and it supports composing an image from several reference images. It replaces the licensed outside image models Meta had used in its apps. On an Arena leaderboard snapshot from 5 July that Meta reported, Muse Image ranks second for text-to-image and edits. Meta also showed Muse Video as an early preview, ranked third for text-to-video, and Meta says it still has gaps in audio sync and fast motion.

Muse Spark 1.1 has a 1 million token context window and improves tool use, computer use, coding and video captioning. The Meta Model API lets outside developers build on Muse, and partners describe it as compatible with OpenAI's API format. Meta has hosted Llama for others before, and this is the first time it sells API access to its own frontier model. TechCrunch reported the price as $1.25 per million input tokens and $4.25 per million output tokens.

xAI releases Grok 4.5, built with Cursor

On 8 July xAI released Grok 4.5, a coding and agent model trained with Cursor, and reports 64.7% on SWE-Bench Pro.

Grok 4.5 is a mixture-of-experts model trained jointly with Cursor on trillions of tokens of Cursor usage data. xAI reports 83.3% on Terminal-Bench 2.1, against 84.3% for Fable at max effort. On DeepSWE 1.0, a benchmark by Datacurve that Artificial Analysis ran with each provider's own harness, Grok 4.5 scored 62.0%.

xAI claims the model uses 4.2 times fewer output tokens than Claude Opus 4.8 at max effort. It costs $2 per million input tokens and $6 per million output tokens for prompts under 200,000 tokens. Snapshots of codebases reportedly leaked into the training data, which makes some of the scores optimistic, and it's not clear which benchmarks are affected or by how much.

Also in the news

  • GPT-Live OpenAI released a full-duplex voice model on 8 July that listens and speaks at the same time and, according to Simon Willison's account, hands search and reasoning to GPT-5.5 while it keeps talking, replacing the GPT-4o-era voice mode for paid users, with a mini version for free users.
  • Anthropic interpretability researchers reported on 6 July a small set of internal "J-space" patterns in Claude, found with a Jacobian-based method, that run silently apart from chain-of-thought text and that Claude can describe and adjust on request, which they compare to global-workspace accounts of conscious access in neuroscience.
  • AlphaEvolve became generally available to all Google Cloud customers on the Gemini Enterprise Agent Platform on 9 July, after a private preview that began in December 2025. Users give it a baseline algorithm and goals, and it evolves optimised code that people can read.
  • Robostral Navigate is Mistral's first embodied model, released on 8 July, an 8 billion parameter model that steers wheeled, legged and flying robots from one RGB camera and language instructions, and Mistral reports 76.6% success on unseen R2R-CE environments.
  • SWE-1.7 from Cognition, released on 8 July, is post-trained from Kimi K2.7, scores 42.3% on FrontierCode 1.1 Main and runs at 1,000 tokens per second on Cerebras.
  • Wan-Dancer-14B from Alibaba's Qwen team, released on 10 July, is an open 14 billion parameter model that turns music into minute-long dance videos at 720p and 30 frames per second.

People

  • Fidji Simo stepped down on 9 July as OpenAI's CEO of Applications, the company's No. 2, and moved to a part-time advisory role after a medical leave that began in April ran longer than expected.