The Week in AI
The week in brief
OpenAI published 722 mathematics manuscripts on 6 October, grouped into 372 families of results, all produced by an internal model it has not released.
OpenAI made three announcements on 6 and 7 October. It published a batch of machine-produced math results with many proofs formalized in Lean, rolled out GPT-6 in ChatGPT with answers that can include charts and small tools, and opened a cheap Decisions API for classification work. Anthropic released Claude Haiku 5.5 on 7 October at $0.10 per million input tokens for shorter requests. On 8 October Google put a planning agent into Gemini Enterprise in private preview.
Two large open-weight models from Western labs were announced on 5 and 6 October, and neither can be downloaded yet. Mistral AI previewed Mistral Large 4 and Reflection AI unveiled Beam, and both say weights follow later in October. Every capability figure in this week's stories is company-reported, and no independent evaluator has published results on any of them so far.
OpenAI published 722 math manuscripts from an unreleased model
OpenAI put 722 mathematics manuscripts, grouped into 372 families of related results, in a public GitHub repository on 6 October, all produced by an internal frontier model it has not released.
What happened
Each family holds a main result plus supporting arguments or alternative proofs. Many of the proofs come with Lean formalizations. Lean is a proof assistant, a language in which a proof is written so that software checks every step, which means those parts can be verified mechanically. OpenAI says more formal proofs will be added.
OpenAI also released 10 detailed summaries of how the model reasoned its way to particular results, along with compute estimates and statistics on the problems it attempted. It estimates the average result used about three hours of ChatGPT Pro thinking compute. Latent Space put the number of problems attempted at about 4,000, and other outlets reported the same figure, which the Atlas didn't find on OpenAI's own page.
The release follows OpenAI's earlier Navier-Stokes announcement, which concerned a single result. This batch covers algebra, number theory, computer science, mathematical logic and topology, according to Ynetnews.
How OpenAI framed the verification
OpenAI consulted the IAS Advisory Group on Mathematics and AI on how to release the batch. It says the results are at different stages of verification and may contain errors. OpenAI reportedly withdrew three of the papers, according to a history file in its GitHub repository, and the records don't say which ones or why.
Reported counts differ slightly. Ynetnews describes more than 700 papers solving 377 previously unsolved problems, and OpenAI's own grouping is 372 result families, and the records don't explain the gap of five. Ynetnews also reports that some results relate to three of the Millennium Prize Problems. "Relate to" is Ynetnews's wording, and no record says any of the three has been solved.
Why it matters
This is the largest batch of machine-produced math results the Atlas has logged from any lab, and it comes with the proofs, the formal checks and a cost estimate in one public place. Mathematicians can read and check the work directly, and the Lean parts can be checked by software without trusting OpenAI.
The model itself stays internal. Outside groups can verify the proofs, but they can't rerun the model on new problems or see how often it failed beyond OpenAI's own statistics. The three hours of Pro thinking per result is the figure to use when someone asks what this costs, and it's OpenAI's estimate.
What we don't know
How many of the 372 families hold up under outside review is unknown. What share of proofs are fully formalized in Lean, and when the additional formal proofs arrive, wasn't in the records. OpenAI hasn't named the model or said whether it will be released.
Mistral previewed Large 4 and Reflection AI unveiled Beam
Mistral AI opened a preview API for the 1 trillion parameter Mistral Large 4 on 6 October, a day after Reflection AI unveiled the 501 billion parameter Beam, and both say open weights will ship later in October.
What happened
Mistral Large 4 is Mistral's largest model so far, up from the 675 billion parameters of Mistral Large 3. It is multimodal and runs as a preview API on Mistral Studio. Mistral says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters, and open weights are due by the end of October.
Both models use a mixture-of-experts design. Each token is routed through a small set of expert subnetworks, so only part of the model does work at any one time, and that keeps inference cost closer to a much smaller model. Mistral's blog gives 1 trillion total and 49 billion active parameters for Large 4. Its documentation lists 1.05 trillion total and 52 billion active, and the company hasn't explained the gap.
Beam is Reflection AI's first model release. It is text-only, aimed at coding and agent tasks, and has 501 billion total and 23 billion active parameters with a 1 million token context window. Reflection says it pretrained Beam on 23.8 trillion tokens and then trained it heavily with reinforcement learning (RL), where the model improves by being scored on tasks it attempts. Beam is a research preview, and the weights and a technical report are due later in October.
What the two labs claim
Mistral says Large 4 significantly outperforms any open-weight model from the US or Europe and is competitive with the strongest open models globally. That second phrase points at the Chinese open-weight models, which Mistral doesn't claim to beat. Mistral didn't publish scores the Atlas could check on 6 October.
Reflection compared Beam with Z.ai's GLM-5.2, which has about 744 billion total and 40 billion active parameters. Reflection says Beam matches GLM-5.2 on reasoning benchmarks while using 3 to 4 times less inference compute. Beam has a little over half of GLM-5.2's active parameters, which doesn't by itself account for a 3 to 4 times saving, so the technical report will have to show where the rest comes from.
Why it matters
Both labs measured themselves against open models from China. For a practitioner choosing an open model at this size, the options have mostly come from Chinese labs, and Large 4 and Beam are the European and US entries for October. StepFun added another Chinese model at this size on 8 October, Step 5 Preview, with 600 billion total and 27 billion active parameters, though it is available only through an API.
Until its weights ship, Mistral is red-teaming a version of Large 4 with reduced moderation together with cybersecurity partners. An open-weight model can't be recalled once released, and Anthropic, OpenAI and Google keep the weights of their top models. Mistral's red-teaming period is how it is setting its cyber position before release.
What we don't know
Which of Mistral's two parameter counts is correct is unknown, and so is whether the released weights will be the moderated version or the one being red-teamed. The benchmark scores behind both companies' comparisons aren't in the records, and Reflection hasn't announced license terms. It's also unknown how Beam compares with GLM-5.3, the newer model in the same family.
OpenAI put GPT-6 into ChatGPT with Intelligent UI
OpenAI rolled out GPT-6 in ChatGPT on 7 October with a feature called Intelligent UI, with paid tiers running on Sol and free users on Luna.
What happened
With Intelligent UI, a ChatGPT answer can mix text with charts, diagrams, forms, tappable buttons and small interactive tools such as a calculator. The model decides which format fits the question. OpenAI built a library of streamable interface components and a compiler that draws the interface while the model is still generating.
The model can also start answering while it keeps thinking and send partial answers. OpenAI says GPT-6 resists attempts to bypass its safety training better than GPT-5.6, and the Atlas has no independent test of that. OpenAI says the model is built for ChatGPT's more than 1.2 billion weekly users.
The Decisions API and Ultrafast
A day earlier, on 6 October, OpenAI opened the Decisions API in public beta. It takes text, images or both and returns one of three typed answers, a probability that a condition is true, a choice from a fixed set, or a score against a rubric. That suits moderation, routing and grading jobs that currently go through a general chat endpoint.
The only supported model is gpt-6-luna, at $0.10 per million input tokens with no charges for output or cached tokens. OpenAI says it runs about 10 times faster than the Responses API and gives no benchmark for that. On 8 October OpenAI also added an Ultrafast service tier for GPT-6.1 Sol in the Responses API, which shortens the time between output tokens and is open to all API users within rate limits.
Why it matters
Free ChatGPT users get GPT-6 Luna in the same week Google moves free Gemini users to Flash Lite only, from 9 October. Anyone comparing the free versions of the two products on stage will be comparing different tiers than a week earlier.
Intelligent UI hands interface choices to the model. A reply can now be something the user taps or fills in, so a demo or screenshot of a GPT-6 answer depends on what the model chose to render for that prompt.
What we don't know
The records name three GPT-6 variants, Sol, Luna and Astra, and don't say how they differ in size or capability. OpenAI published no benchmark scores for GPT-6 in ChatGPT in the records the Atlas holds. Usage limits per tier and the Ultrafast price weren't in the sources.
Anthropic released Claude Haiku 5.5 with adaptive thinking
Anthropic released Claude Haiku 5.5 on 7 October, its smallest model, with a 1 million token context window at $0.10 per million input tokens and $0.50 per million output tokens for requests up to 100,000 tokens.
What happened
Haiku 5.5 succeeds Haiku 4.5. It has a 1 million token context window and can write up to 128,000 tokens in one response. Anthropic says it is its fastest model, aimed at summaries, classification, browser use and subagent coding work alongside Opus 5.5 and Sonnet 5.5.
It is the first Haiku with adaptive thinking. The developer picks an effort setting, and the model decides how much to reason within it, so easy requests stay quick. Requests longer than 100,000 tokens cost $0.50 per million input tokens and $2.50 per million output tokens. Anthropic says the new model is about 75% cheaper than Haiku 4.5 on average.
What Anthropic reports
All of these scores are Anthropic's. Haiku 5.5 scores 39.2% on Terminal-Bench 4.0, where Haiku 4.5 scored 0.0%. On the offline subset of OSWorld 2.1, a test of operating a computer through its interface, it scores 72.4% against 15.7% for Haiku 4.5.
Anthropic reports 45.9% on Humanity's Last Exam without tools, up from 10.2%. On GDPval-AA v2.1 it reports a score of 1620, against 735 for Haiku 4.5.
Anthropic also halved cache read prices for Sonnet 5.5, which launched on 28 September. It added a monthly API credit for Max and Team subscribers.
Why it matters
The jump on Terminal-Bench and OSWorld is what makes the subagent use case plausible. A small model that scored zero on terminal tasks couldn't take delegated coding work, and at 39.2% Anthropic is pitching it for exactly that.
The price lines up with OpenAI's Decisions API, which also charges $0.10 per million input tokens and launched the day before. Classification and routing at very low cost now has a model from each lab, with OpenAI charging nothing for output and Anthropic charging $0.50 per million output tokens.
What we don't know
No independent evaluator has published results for Haiku 5.5 yet. The records don't say how Anthropic calculated the 75% average saving or which usage mix it assumed.
Google previewed a Gemini Enterprise agent that delegates to subagents
Google announced a universal agent in the Gemini Enterprise app on 8 October, in private preview for enterprise customers, that takes an objective, plans the work and hands parts of it to subagents.
What happened
Until now, Gemini Enterprise was a place for companies to build and govern their own agents. The new agent takes a goal instead of step-by-step instructions, plans the work, and delegates pieces to job-specific subagents that use custom skills and tools. It picks a model automatically, and users can choose another one, including Anthropic's Claude.
The agent has its own Workspace account with an email address, so its actions appear in the audit trail under the agent's name. It connects to Workspace, Microsoft 365, Slack, Jira, Snowflake, BigQuery and any server that speaks the Model Context Protocol (MCP), and users follow its progress in a tasks inbox.
Sundar Pichai, Google's chief executive, said at the event that Gemini has over 1 billion monthly active users and that nearly 90% of Fortune 100 companies use Gemini Enterprise. Both figures are Google's own.
Why it matters
Giving the agent its own account is a concrete answer to a question companies ask about agents, which is who did what. Actions are logged under the agent's identity, so an administrator can separate them from the employee who set the goal.
Google letting customers pick Claude inside its own product is a case of a lab routing work to a rival's model.
What we don't know
Google gave no price, no general availability date and no count of preview customers in the records. Which model the agent picks by default, and how it decides, wasn't described.
Also worth knowing
Security and provenance
- Anthropic expanded its Cyber Verification Program on 6 October and folded Project Glasswing into it, with three access tiers for vetted security professionals, six days after Google gave Gemini 4 Argon to cyber defenders first.
- Glasswing figures reported by Reuters, as summarized by Techmeme, have Anthropic finding more than 5,500 verified vulnerabilities from April to October and partners finding more than 129,000 from April to July, and these are company-reported counts.
- Critical Infrastructure Defense Program is Anthropic's second security program of the week, launched on 8 October with 11 founding partners including CrowdStrike, Dragos, Palo Alto Networks and Rockwell Automation, giving defenders of grids, water utilities, factories and transport networks frontier Claude models, threat research and on-site engineers.
- Goodfire launched monitors for Baseten customers on 8 October that read a model's internal activations with small classifiers called probes, and it reports catching 94% of malicious hacking sessions while monitoring about 1,500 Kimi K3 sessions for about $51, against about $10,000 for a top-tier model checking every step.
- SynthID Detector is a public Google site opened on 7 October that checks images, video and audio for SynthID watermarks from Google, OpenAI, Nvidia and Kakao, with about 10 checks per user per day and Apple support promised later.
- textGrain is OpenAI's statistical text watermark, announced on 5 October for the EU AI Act, and OpenAI reports its detector catches about 95% of 400-token passages at a 1% false positive rate but only 17% once a quarter of the words are replaced.
Open and on-device models
- EmbeddingGemma 2 is an open 740 million parameter embedding model from Google DeepMind, released on 6 October and built on Gemma 4, that maps text, code, images, video and audio into one 768-dimensional space so a text query can be compared directly with a video clip.
- EmbeddingGemma 2's encoders are modular, and Google reports about 191MB of active RAM for the 270 million parameter text-only part on a Pixel 11 Pro.
- Google AI Edge Foresight is a free experimental Mac app released on 8 October that transcribes meetings and writes notes offline with on-device models including EmbeddingGemma 2.
- pplx-embed-v2-late is a pair of open-weight multimodal retrieval models from Perplexity at 0.6 billion and 9 billion parameters under the MIT license, and Perplexity reports 92.4% on MADQA for the larger one.
- DeepSeek V4 Flash will run locally on RTX Spark PCs quantized to 1.6 bits, Microsoft said at its 7 October Windows keynote, alongside a Nvidia Nemotron of more than 70 billion parameters and llama.cpp support in Windows ML.
- Falcon-Emirati-7B was announced by the Technology Innovation Institute, a model built on Falcon-H1-Arabic and specialised in the Emirati Arabic dialect and culture.
Products and pricing
- Gemini's free tier drops to Flash Lite only from 9 October, standard Flash moves to the $4.99 per month Google AI Plus plan, and Gemini Pro and Deep Think stay only on Google AI Pro at $19.99 and Ultra at $99.99 per month.
- ChatGPT image ads were announced by OpenAI on 5 October, with labeled ads next to image generation results for Free and Go users and US testing later in October.
- Nano Banana 2.1 is Google DeepMind's new Flash-tier image model, listed on OpenRouter at $1.50 per million input tokens, $7.50 per million output tokens and $30 per million image output tokens.
- Step 5 Preview from StepFun is on OpenRouter at $1 per million input tokens and $2.70 per million output tokens.
- Liquid Inference from Architect Financial Technologies auctions each request among providers serving the named model and charges the lowest offer that meets the buyer's limits on price, latency and region.
- Playground is a Google Labs service launched in the US on 7 October that builds playable browser games from text prompts.
Agents
- Personal Agent Protocol is an open standard published on 6 October by Meta with Walmart, Stripe, Sierra and others to help websites tell bots acting for a user from malicious ones, after Amazon began blocking Meta's Muse agent from its retail site.
- Ironclad worked with OpenAI on 11 contracting tasks, and OpenAI reports GPT-6 Astra scored 55.0% against 41.6% for GPT-5.6 Sol, on a page with no publication date.
Hardware
- RTX Spark is an Arm-based Windows platform from NVIDIA and Microsoft announced on 7 October with up to 128GB of unified memory, and The Verge reports the Surface Laptop Ultra built on it starts at $2,599 and ships on 16 October.
Research
- Vals AI reports that a team of Claude Opus 5.5 agents found two candidate room-temperature antiferromagnetic semiconductors for computer memory, predicted by calculation and not yet tested in a lab.
Money
- OpenAI told investors its annualized revenue was nearing $50 billion at the end of September, according to the Financial Times, which is well below the $70 billion reported earlier.
- DeepSeek is reportedly close to a funding round of more than $11 billion.
- Kling, Kuaishou's video unit, has reportedly picked banks for a Hong Kong IPO targeting at least $1 billion.
- Arena, which runs the model leaderboard, reportedly raised $200 million at a $3.1 billion valuation and launched an Alignment Index.
- Nous Research reportedly raised $90 million at a $1.5 billion valuation and launched AI agents for business users.
People this week
- Rohit Prasad, who helped build Alexa and led Amazon's AGI organization and the Nova models, will become CEO of Boston Dynamics, the Hyundai Motor Group robotics company, to push intelligent robots into commercial use.
- Luo Fuli, who leads Xiaomi's MiMo model team, was reportedly promoted to vice president level, according to 36Kr.
What did not change
Every capability number in this week's stories is company-reported. No independent evaluator has published results for Mistral Large 4, Beam, GPT-6 or Haiku 5.5, and OpenAI's math results are at different stages of verification by OpenAI's own account.
The most cyber-capable models from Anthropic, OpenAI and Google still go to vetted defenders before anyone else. Anthropic's 6 October expansion added tiers to that arrangement, and its 8 October infrastructure program extends it through security firms and equipment makers.
Neither of the week's large open-weight models from Western labs can be downloaded yet. Until Mistral and Reflection ship weights, the strongest downloadable models at this size are still the Chinese ones they compare against.
What to watch next
- On 9 October Google's Gemini free tier drops to Flash Lite only, and AI Plus loses Gemini Pro.
- On 15 October the Nvidia Nemotron model of more than 70 billion parameters for local use on Windows is expected, according to Microsoft's keynote.
- On 16 October the Surface Laptop Ultra on RTX Spark ships, with the $5,999 Surface RTX Spark Dev Box following in November.
- Later in October Reflection AI says it will release Beam's weights and a technical report, which should show the scores behind its GLM-5.2 comparison.
- By the end of October Mistral says Large 4's weights will ship, and the release will show which moderation settings the public version carries.
- Later in October OpenAI begins US testing of image ads in ChatGPT for Free and Go users.
- OpenAI expects the Decisions API to reach general availability in the coming weeks and has given no date.
- OpenAI says more Lean proofs will be added to the math repository and has given no date.
- Google has given no date for general availability of the Gemini Enterprise agent.