The week in AI
The week in brief
OpenAI published 722 mathematics manuscripts on 6 October, grouped into 372 families of results, all produced by an internal model it has not released.
OpenAI made three announcements on 6 and 7 October. It published a batch of machine-produced math results with many proofs formalized in Lean, rolled out GPT-6 in ChatGPT with answers that can include charts and small tools, and opened a cheap Decisions API for classification work. Anthropic released Claude Haiku 5.5 on 7 October at $0.10 per million input tokens for shorter requests.
Two large open-weight models from Western labs were announced on 5 and 6 October, and neither can be downloaded yet. Mistral AI previewed Mistral Large 4 and Reflection AI unveiled Beam, and both say weights follow later in October. Every capability figure in this week's stories is company-reported, and no independent evaluator has published results on any of them so far.
OpenAI published 722 math manuscripts from an unreleased model
OpenAI put 722 mathematics manuscripts, grouped into 372 families of related results, in a public GitHub repository on 6 October, all produced by an internal frontier model it has not released.
What happened
Each family holds a main result plus supporting arguments or alternative proofs. Many of the proofs come with Lean formalizations. Lean is a proof assistant, a language in which a proof is written so that software checks every step, which means those parts can be verified mechanically. OpenAI says more formal proofs will be added.
OpenAI also released 10 detailed summaries of how the model reasoned its way to particular results, along with compute estimates and statistics on the problems it attempted. It estimates the average result used about three hours of ChatGPT Pro thinking compute. Latent Space put the number of problems attempted at about 4,000, and other outlets reported the same figure, which the atlas didn't find on OpenAI's own page.
The release follows OpenAI's earlier Navier-Stokes announcement, which concerned a single result. This batch covers algebra, number theory, computer science, mathematical logic and topology, according to Ynetnews.
How OpenAI framed the verification
OpenAI consulted the IAS Advisory Group on Mathematics and AI on how to release the batch. It says the results are at different stages of verification and may contain errors. The Lean formalizations cover many of the proofs, and the records don't give the share.
Reported counts differ slightly. Ynetnews describes more than 700 papers solving 377 previously unsolved problems, and OpenAI's own grouping is 372 result families, and the records don't explain the gap of five. Ynetnews also reports that some results relate to three of the Millennium Prize Problems. "Relate to" is Ynetnews's wording, and no record says any of the three has been solved.
Why it matters
This is the largest batch of machine-produced math results the atlas has logged from any lab, and it comes with the proofs, the formal checks and a cost estimate in one public place. Mathematicians can read and check the work directly, and the Lean parts can be checked by software without trusting OpenAI.
The model itself stays internal. Outside groups can verify the proofs, and they can't rerun the model on new problems or see how often it failed beyond OpenAI's own statistics. The three hours of Pro thinking per result is the figure to use when someone asks what this costs, and it's OpenAI's estimate.
What we don't know
How many of the 372 families hold up under outside review is unknown. What share of proofs are fully formalized in Lean, and when the additional formal proofs arrive, wasn't in the records. OpenAI hasn't named the model or said whether it will be released.
Mistral previewed the 1 trillion parameter Mistral Large 4
Mistral AI released a preview API for Mistral Large 4 on 6 October, a 1 trillion parameter model with 49 billion parameters active per token, and says open weights will follow by the end of October.
What happened
Mistral Large 4 is Mistral's largest model so far, up from the 675 billion parameters of Mistral Large 3. It is multimodal, and it runs as a preview API on Mistral Studio. Mistral says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters.
Large 4 is a mixture-of-experts model. Each token is routed through a small set of expert subnetworks, so only part of the model does work at any one time, and that keeps inference cost closer to a much smaller model. Mistral's blog gives 1 trillion total and 49 billion active parameters. Mistral's documentation lists 1.05 trillion total and 52 billion active, and the company hasn't explained the gap.
What Mistral claims
Mistral says Large 4 significantly outperforms any open-weight model from the US or Europe and is competitive with the strongest open models globally. That second phrase points at the Chinese open-weight models, which Mistral doesn't claim to beat. The atlas has no independent benchmark result for Large 4, and Mistral didn't publish scores the atlas could check on 6 October.
Why it matters
Until the weights ship, Mistral is red-teaming a version with reduced moderation together with cybersecurity partners. That comes after a result logged on 29 September, when Anthropic reported that Zhipu's open-weights GLM-5.3 built working exploits about two-thirds as often as its own gated Mythos Preview. An open-weight model can't be recalled once released.
Mistral's approach is a short closed window before an open release. Anthropic, OpenAI and Google gate their top models for longer and keep the weights. A lab releasing 1 trillion parameters of open weights has to decide its cyber position before release, and Mistral's red-teaming period is how it is doing that.
What we don't know
Which of Mistral's two parameter counts is correct is unknown. The scores behind "competitive with the strongest open models globally" weren't available to the atlas. Whether the released weights will be the moderated version or the one being red-teamed with reduced moderation is also unknown.
Reflection AI unveiled Beam, a 501 billion parameter model
Reflection AI unveiled Beam on 5 October, a text-only mixture-of-experts model with 501 billion total and 23 billion active parameters, and says it matches Z.ai's GLM-5.2 on reasoning benchmarks with 3 to 4 times less inference compute.
What happened
Beam is Reflection AI's first model release. It is text-only, aimed at coding and agent tasks, and has a 1 million token context window. Reflection says it pretrained Beam on 23.8 trillion tokens and then trained it heavily with reinforcement learning (RL), where the model improves by being scored on tasks it attempts.
Beam is available as a research preview. Reflection says the weights and a technical report will come later in October.
The comparison with GLM-5.2
Reflection chose a Chinese open model as its benchmark. GLM-5.2 has about 744 billion total and 40 billion active parameters, so Beam is smaller on both counts. Reflection says Beam matches GLM-5.2 on reasoning benchmarks while using 3 to 4 times less inference compute, and nobody has checked that claim independently.
Active parameters drive most of the per-token cost in a mixture-of-experts model. Beam has a little over half of GLM-5.2's active parameters, which doesn't by itself account for a 3 to 4 times saving. The technical report will have to show where the rest comes from.
Why it matters
Mistral and Reflection announced large open-weight models within two days of each other, and both compared themselves against open models from China. The GLM line is the same family Anthropic tested for exploit ability on 29 September with GLM-5.3. For a practitioner choosing an open model at this size, the options have mostly come from Chinese labs, and Beam and Large 4 are the US and European entries for October.
What we don't know
The benchmark scores behind "matches GLM-5.2" aren't in the records. The license terms for the weights haven't been announced. It's also unknown how Beam compares with GLM-5.3, the newer model in the same family.
OpenAI put GPT-6 into ChatGPT with Intelligent UI
OpenAI rolled out GPT-6 in ChatGPT on 7 October with a feature called Intelligent UI, with paid tiers running on Sol and free users on Luna.
What happened
With Intelligent UI, a ChatGPT answer can mix text with charts, diagrams, forms, tappable buttons and small interactive tools such as a calculator. The model decides which format fits the question. OpenAI built a library of streamable interface components and a compiler that draws the interface while the model is still generating.
The model can also start answering while it keeps thinking and send partial answers. OpenAI says GPT-6 resists attempts to bypass its safety training better than GPT-5.6, and the atlas has no independent test of that. OpenAI says the model is built for ChatGPT's more than 1.2 billion weekly users.
The Decisions API
A day earlier, on 6 October, OpenAI opened the Decisions API in public beta. It takes text, images or both and returns one of three typed answers, a probability that a condition is true, a choice from a fixed set, or a score against a rubric. That suits moderation, routing and grading jobs that currently go through a general chat endpoint.
The only supported model is gpt-6-luna, at $0.10 per million input tokens with no charges for output or cached tokens. OpenAI says it runs about 10 times faster than the Responses API and gives no benchmark for that. It expects general availability in the coming weeks.
Why it matters
Free ChatGPT users get GPT-6 Luna in the same week Google moves free Gemini users to Flash Lite only, from 9 October. Anyone comparing the free versions of the two products on stage will be comparing different tiers than a week earlier.
Intelligent UI hands interface choices to the model. A reply can now be something the user taps or fills in, so a demo or screenshot of a GPT-6 answer depends on what the model chose to render for that prompt.
What we don't know
The records name three GPT-6 variants, Sol, Luna and Astra, and don't say how they differ in size or capability. OpenAI published no benchmark scores for GPT-6 in ChatGPT in the records the atlas holds. Usage limits per tier weren't in the sources.
Anthropic released Claude Haiku 5.5 with adaptive thinking
Anthropic released Claude Haiku 5.5 on 7 October, its smallest model, with a 1 million token context window at $0.10 per million input tokens and $0.50 per million output tokens for requests up to 100,000 tokens.
What happened
Haiku 5.5 succeeds Haiku 4.5. It has a 1 million token context window and can write up to 128,000 tokens in one response. Anthropic says it is its fastest model, aimed at summaries, classification, browser use and subagent coding work alongside Opus 5.5 and Sonnet 5.5.
It is the first Haiku with adaptive thinking. The developer picks an effort setting, and the model decides how much to reason within it, so easy requests stay quick. Requests longer than 100,000 tokens cost $0.50 per million input tokens and $2.50 per million output tokens. Anthropic says the new model is about 75% cheaper than Haiku 4.5 on average.
What Anthropic reports
All of these scores are Anthropic's. Haiku 5.5 scores 39.2% on Terminal-Bench 4.0, where Haiku 4.5 scored 0.0%. On the offline subset of OSWorld 2.1, a test of operating a computer through its interface, it scores 72.4% against 15.7% for Haiku 4.5.
Anthropic reports 45.9% on Humanity's Last Exam without tools, up from 10.2%. On GDPval-AA v2.1 it reports a score of 1620, against 735 for Haiku 4.5.
Anthropic also halved cache read prices for Sonnet 5.5, which launched on 28 September. It added a monthly API credit for Max and Team subscribers.
Why it matters
The jump on Terminal-Bench and OSWorld is what makes the subagent use case plausible. A small model that scored zero on terminal tasks couldn't take delegated coding work, and at 39.2% Anthropic is pitching it for exactly that.
The price lines up with OpenAI's Decisions API, which also charges $0.10 per million input tokens and launched the day before. Classification and routing at very low cost now has a model from each lab, with OpenAI charging nothing for output and Anthropic charging $0.50 per million output tokens.
What we don't know
No independent evaluator has published results for Haiku 5.5 yet. The records don't say how Anthropic calculated the 75% average saving or which usage mix it assumed.
Also worth knowing
Security and provenance
- Anthropic expanded its Cyber Verification Program on 6 October and folded Project Glasswing into it, with three access tiers for vetted security professionals, six days after Google gave Gemini 4 Argon to cyber defenders first.
- Glasswing figures reported by Reuters, as summarized by Techmeme, have Anthropic finding more than 5,500 verified vulnerabilities from April to October and partners finding more than 129,000 from April to July, and these are company-reported counts.
- Fewer safeguards for vetted users is reported in the atlas's record of the Anthropic program, and it didn't appear on any evidence page the atlas checked.
- SynthID Detector is a public Google site opened on 7 October that checks images, video and audio for SynthID watermarks from Google, OpenAI, Nvidia and Kakao, with about 10 checks per user per day and Apple support promised later.
Open models
- EmbeddingGemma 2 is an open 740 million parameter embedding model from Google DeepMind, released on 6 October and built on Gemma 4, that maps text, code, images, video and audio into one 768-dimensional space so a text query can be compared directly with a video clip.
- EmbeddingGemma 2's encoders are modular, and Google reports about 191MB of active RAM for the 270 million parameter text-only part on a Pixel 11 Pro.
- Falcon-Emirati-7B was announced by the Technology Innovation Institute, a model built on Falcon-H1-Arabic and specialised in the Emirati Arabic dialect and culture.
Products and pricing
- Gemini's free tier drops to Flash Lite only from 9 October, standard Flash moves to the $4.99 per month Google AI Plus plan, and Gemini Pro and Deep Think stay only on Google AI Pro at $19.99 and Ultra at $99.99 per month.
- ChatGPT image ads were announced by OpenAI on 5 October, with labeled ads next to image generation results for Free and Go users and US testing later in October.
- Nano Banana 2.1 is Google DeepMind's new Flash-tier image model, listed on OpenRouter at $1.50 per million input tokens, $7.50 per million output tokens and $30 per million image output tokens.
- textGrain is OpenAI's new statistical text watermark, announced on 5 October, which will be applied automatically to ChatGPT and Codex text in the EU over the coming weeks, with an opt-in for API customers worldwide.
Agents
- Personal Agent Protocol is an open standard published on 6 October by Meta with Walmart, Stripe, Sierra and others to help websites tell bots acting for a user from malicious ones, after Amazon began blocking Meta's Muse agent from its retail site.
- Ironclad worked with OpenAI on 11 contracting tasks, and OpenAI reports GPT-6 Astra scored 55.0% against 41.6% for GPT-5.6 Sol, on a page with no publication date.
- Grok Bot will route some tasks to rival models including Claude Opus 5.5, Elon Musk said, according to The Information.
Hardware
- RTX Spark is an Arm-based Windows platform from NVIDIA and Microsoft announced on 7 October with up to 128GB of unified memory, and The Verge reports the Surface Laptop Ultra built on it starts at $2,599 and ships on 16 October.
Research
- Vals AI reports that a team of Claude Opus 5.5 agents found two candidate room-temperature antiferromagnetic semiconductors for computer memory, predicted by calculation and not yet tested in a lab.
- Nvidia published a write-up on fine-tuning one Nemotron model family to gold-level results at both the International Olympiad in Informatics and the International Mathematical Olympiad.
Money
- DeepSeek is reportedly close to a $12 billion funding round backed by Tencent.
- Kling, Kuaishou's video unit, has reportedly picked banks for a Hong Kong IPO worth more than $1 billion, according to The Information.
- Nous Research reportedly raised a $90 million Series B at a $1.5 billion valuation to scale its Hermes Agent.
- Vinci, an engineering AI startup, reportedly raised $250 million at a $1.5 billion valuation.
People this week
- Rohit Prasad, who helped build Alexa and led Amazon's AGI organization and the Nova models, will become CEO of Boston Dynamics, the Hyundai Motor Group robotics company, to push intelligent robots into commercial use.
What did not change
Every capability number in this week's stories is company-reported. No independent evaluator has published results for Mistral Large 4, Beam, GPT-6 or Haiku 5.5, and OpenAI's math results are at different stages of verification by OpenAI's own account.
The most cyber-capable models from Anthropic, OpenAI and Google still go to vetted defenders before anyone else. Anthropic's 6 October expansion adds tiers to that arrangement.
Neither of the week's large open-weight models can be downloaded yet. Until Mistral and Reflection ship weights, the strongest downloadable models at this size are still the Chinese ones they compare against.
What to watch next
- On 9 October Google's Gemini free tier drops to Flash Lite only, and AI Plus loses Gemini Pro.
- On 16 October the Surface Laptop Ultra on RTX Spark ships, with the $5,999 Surface RTX Spark Dev Box following in November.
- Later in October Reflection AI says it will release Beam's weights and a technical report, which should show the scores behind its GLM-5.2 comparison.
- By the end of October Mistral says Large 4's weights will ship, and the release will show which moderation settings the public version carries.
- Later in October OpenAI begins US testing of image ads in ChatGPT for Free and Go users.
- OpenAI expects the Decisions API to reach general availability in the coming weeks and has given no date.
- OpenAI says more Lean proofs will be added to the math repository and has given no date.
- Anthropic has given no date for publishing the details of its three Cyber Verification Program tiers.