The week in AI

Week of 2 Feb to 8 Feb 2026

The week in brief

Anthropic released Claude Opus 4.6 and OpenAI released GPT-5.3-Codex on 5 February, and OpenAI reports its model scores 77.3% on Terminal-Bench 2.0.

The two releases landed on the same day and both aimed at long-running coding and agent work. Anthropic's Opus 4.6 is the first Opus with a 1M-token context window, offered in beta. OpenAI's GPT-5.3-Codex merges its coding and general knowledge-work models into one. OpenAI also shipped a Codex desktop app for macOS on 2 February.

Anthropic published three engineering and research posts alongside its model. One describes a C compiler written by 16 parallel Claude agents. One measures how much server setup moves agentic coding scores. One reports more than 500 high-severity open-source vulnerabilities found by Claude. In open weights, Alibaba released Qwen3-Coder-Next and Shanghai AI Laboratory released the trillion-parameter Intern-S1-Pro.

Anthropic released Claude Opus 4.6 with a 1M-token context

Claude Opus 4.6 adds a 1M-token context window in beta and 128K output tokens, and keeps its price at $5 per million input tokens and $25 per million output tokens.

Above 200K tokens of input the price rises to $10 per million input tokens and $37.50 per million output tokens. The release adds adaptive thinking, where the model decides how much to reason on each request within one of four effort levels you set. It also adds context compaction in beta, which summarizes older parts of a long session so an agent can keep working past the window. Claude Code gained agent teams as a research preview, so several Claude instances can split up one task. A faster serving mode followed as a research preview on 7 February.

Anthropic reports 76% on MRCR v2 with 8 needles at 1M tokens, a test of finding eight planted items in a very long document, against 18.5% for Sonnet 4.5. On GDPval-AA, an Artificial Analysis benchmark of economically useful work tasks, Anthropic reports Opus 4.6 is 144 Elo points above GPT-5.2 and 190 above Opus 4.5. Anthropic also cites 80.8% on SWE-bench Verified and 65.4% on Terminal-Bench 2.0.

On the same day Anthropic reported that Opus 4.6 found and validated more than 500 high-severity vulnerabilities in well-tested open-source software, with no custom scaffolding. Anthropic says the model worked like a human researcher, reading past fixes and looking for similar bug patterns, and found bugs in projects that had years of fuzzing behind them. Humans validated each one before reporting, Anthropic is patching them with maintainers, and it described the misuse safeguards it added.

OpenAI released GPT-5.3-Codex and a Codex app for macOS

OpenAI reports GPT-5.3-Codex scores 77.3% on Terminal-Bench 2.0, up from 64.0% for GPT-5.2-Codex.

GPT-5.3-Codex combines the coding skills of GPT-5.2-Codex with the knowledge work of GPT-5.2 in one model that OpenAI says runs 25% faster. OpenAI reports OSWorld-Verified, a test of operating a desktop computer, rising from 38.2% to 64.7%. SWE-Bench Pro barely moved, from 56.4% to 56.8%. OpenAI calls it the first model "instrumental in creating itself" because early versions helped debug its own training and manage its deployment.

It is also the first model OpenAI classifies as High for cybersecurity under its own framework. OpenAI paired the release with $10 million in API credits for defenders and Trusted Access for Cyber, an identity-verified program that gives vetted security teams easier access to its more cyber-capable models. There was no API access during the week.

Three days earlier, on 2 February, OpenAI released the Codex app for macOS. It runs many coding agents in parallel threads, gives each one its own git worktree so their changes do not collide, and adds Skills and scheduled Automations. OpenAI says Codex usage doubled after GPT-5.2-Codex and more than a million developers used it in the prior month. Free and Go users got Codex for a limited time, and paid plan limits were doubled.

The two labs' Terminal-Bench 2.0 figures sit 12 points apart, 77.3% for GPT-5.3-Codex and 65.4% for Opus 4.6, both company-reported. Anthropic's own infrastructure paper from the same day, covered below, shows server configuration alone can move that benchmark by 6 points.

Anthropic published two experiments on agentic coding

Anthropic reports that 16 parallel Claude agents wrote a C compiler of about 100,000 lines of Rust in two weeks that boots Linux 6.9.

The compiler run took about 2,000 Claude Code sessions and roughly $20,000 of API use. It boots Linux 6.9 on x86, ARM and RISC-V, passes 99% of compiler test suites including the GCC torture tests, and builds QEMU, FFmpeg, SQLite, PostgreSQL and Redis. It has no 16-bit x86 backend, relies on outside tools for assembling and linking, and its output is slower than GCC with optimizations turned off.

The second post measured noise in agentic coding evaluations. With strict container resource limits, 5.8% of Terminal-Bench 2.0 tasks failed for infrastructure reasons, against 0.5% with no cap. The best and worst resource setups differed by 6 percentage points, a gap Anthropic found statistically significant. Anthropic's conclusion is that leaderboard margins of a few points can come from how the test machines are configured.

Alibaba and Shanghai AI Laboratory released large open-weight models

Alibaba released Qwen3-Coder-Next, an open 80B coding model with 3B active parameters that it says scores over 70% on SWE-bench Verified.

Qwen3-Coder-Next came out on 4 February under Apache 2.0. It is a mixture-of-experts model, so only 3B of its 80B parameters run for each token. Alibaba trained it from Qwen3-Next-80B-A3B-Base with reinforcement learning (RL) on about 800,000 coding tasks whose results can be checked by running them. It has a native context of 262K tokens and works inside Claude Code, Qwen Code and Cline. The over 70% figure uses the SWE-Agent scaffold and is Alibaba's own.

Shanghai AI Laboratory released Intern-S1-Pro on 2 February under Apache 2.0, describing it as the first one-trillion-parameter scientific multimodal model. It has 512 experts with 8 active, for 22B active parameters, and the lab says it covers over 100 scientific tasks. The lab ran trillion-scale RL with its own XTuner and LMDeploy tools and added a position encoding designed for physical signals.

Also in the news

  • MiniCPM-o 4.5 ModelBest released a 9B model on 4 February that streams vision and speech in both directions at once, and it scores 77.6 on OpenCompass vision, above GPT-4o in ModelBest's comparison.
  • Voxtral Transcribe 2 Mistral AI released Voxtral Realtime under Apache 2.0 on 4 February with streaming transcription at under 200 milliseconds of configurable delay, and Mistral reports about 4% word error rate on FLEURS for Mini Transcribe V2.
  • OpenAI Frontier OpenAI launched an enterprise platform on 5 February for building, deploying and managing agents with shared context, onboarding, feedback and permissions.