Reasoning models across the frontier labs, two years after o1
Two years after o1, reasoning is a default, adaptive capability of every frontier model, scaled by RL and bought by effort level, and the argument has moved on from whether models can reason to cost,…
- Date
- 4 October 2026
- Who
- all frontier labs
- Confidence
- Medium (many 2026 items are one to three weeks old; items marked "reported" rest
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Landmark · Significance: 4/5 · Org(s): all frontier labs · Confidence: Medium (many 2026 items are one to three weeks old; items marked "reported" rest on secondary sources) Primary sources: Claude Platform release notes · OpenAI reasoning guide · Gemini thinking docs · DeepSeek updates log · GPT-6 Astra system card · METR time horizons · METR note · Epoch
One-liner. Two years after o1, reasoning is a default, adaptive capability of every frontier model, scaled by RL and bought by effort level, and the argument has moved on from whether models can reason to cost, legibility, long-horizon reliability and who may use the strongest models.
Who has what (as of 2026-10-04).
| Lab | Latest reasoning-capable flagships | How thinking is controlled | Notes |
|---|---|---|---|
| OpenAI | GPT-6 Astra (staged release 2026-09-03/04), GPT-6 Sol and Luna (2026-09-22, per OpenAI's API changelog), GPT-6.1 Sol (2026-09-29, DevDay; $2 / $10 per million tokens; reported to nearly match Astra at about a fifth of the cost) | Documented effort values none through max (seven), summaries optional, reasoning state preservable across turns; some models reject none (docs) | Astra is the first "Critical" cyber model and has reduced CoT monitorability (B08-50); Unite.ai on GPT-6.1 Sol |
| Anthropic | Fable 5.1 and Mythos 5.1 (2026-09-01), Opus 5.5 (2026-09-22, $4 / $20), Sonnet 5.5 (2026-09-28) | Always-on adaptive thinking on Fable/Mythos and Opus 5.5; effort low to max including xhigh; manual token budgets removed; thinking display "omitted" by default since Mythos Preview and Opus 4.7, and the raw chain of thought is never returned (release notes) | Mythos-class models gated by safeguards (B08-45) |
| Google DeepMind | Gemini 3.1 Pro (2026-02-19); Gemini 3 Deep Think upgrade (2026-02-12); Gemini 3.5 Flash and 3.8 Flash (listed in Google's docs; dates reported); Gemini 4 Argon announced 2026-09-30 / 10-01 (sources differ), limited to trusted cyber defenders at first | thinking_level minimal to high; medium is the default on 3.5 to 3.8 Flash and high on 3.1 Pro; thought summaries and encrypted signatures (docs) | Google says it monitors Argon's CoT and actions and can stop execution (The Hacker News) |
| DeepSeek | V4-Pro and V4-Flash (preview 2026-04-24; official releases 2026-08-13 and 2026-07-31), V4.1-Flash (2026-09-10) | Non-think / Think High / Think Max (API low / high / max); R-series endpoints discontinued 2026-07-24 | MIT licence (B08-47) |
| Moonshot | Kimi K3 (announced 2026-07-16; weights promised by 07-27, posting date unconfirmed) | Reasoning always on, max only at launch (the model card now lists low, high and max) | 2.8T parameters, 104B active, custom licence (B08-49) |
| Meta | Muse Spark (2026-04-08), 1.1 (07-09), 1.2 (08-05), 1.3 (09-02), all proprietary; open-weight Muse Glimmer 30B (2026-08-10) | Instant / Thinking / Contemplating | B08-46 |
| Alibaba, Zhipu, MiniMax, xAI | Qwen3.8-Max (GA 2026-08-03, reported) and a Qwen 4 preview (2026-09-22, reported); GLM-5.2 (2026-06-16) and 5.3 (2026-08-18) per Z.ai's release notes; MiniMax M3 (2026-06-01, reported); Grok 4.5 (July 2026) and 4.7 (September 2026) per xAI's docs, which give the month only and now brand them "SpaceXAI" | Hybrid or always-on thinking, vendor-specific | GLM dates and Grok months checked on vendor pages; the Qwen 4 preview and MiniMax M3 remain single-source aggregator reports, so confidence is Low-Medium (B08-43c) |
Eight things to know.
- Effort settings gave way to automatic adaptive thinking. The change ran from
reasoning_effortin 2024-12 to always-on adaptive thinking in 2026 (B08-06). Open-weight Qwen3's hybrid-then-separate reversal and DeepSeek V4's merge show the same convergence (B08-20, B08-47). - RL is a main training lever alongside pretraining. OpenAI said o3 used about 10x o1's training compute, which Epoch reads as RL compute (B08-24); pretraining kept scaling too (V4-Pro is 1.6T parameters trained on 32T+ tokens, Kimi K3 is 2.8T, Qwen3.8 is 2.4T). DeepSeek's RL went from about $294K to over 10% of pre-training cost (B08-43); ScaleRL found predictable scaling curves (B08-40). Whether it adds capability or only sharpens it is still contested (B08-25).
- Contest benchmarks are saturated, and ARC-AGI-2 was followed by a harder ARC-AGI-3. The evidence is company-announced perfect IMO scores and several 42/42 runs in a third-party harness (B08-48), ARC-AGI-2 at a verified 77.1% for Gemini 3.1 Pro and a reported 84.6% for Deep Think (B08-42, B08-43b), then ARC-AGI-3 at under 1% on launch day and 62.7% to 99.9% (harness-dependent) by September (B08-44).
- Models are now compared by how long a task they can complete. METR's time-horizon work (B08-21a) reported a doubling every 6 to 7 months and a 50% horizon of about 4 hours 49 minutes for Claude Opus 4.5 in January 2026 (95% confidence interval 1h49m to 20h25m) (METR note). Its time-horizons page, updated 2026-05-08, adds Opus 4.6, GPT-5.4, Gemini 3.1 Pro and Claude Mythos Preview (early) and says measurements above 16 hours are unreliable with the current suite (METR); as fetched it does not list Opus 5.5, GPT-5.6 Sol, Fable 5 or GPT-6, so this chapter makes no claim about their horizons. See B10.
- Labs have been showing less of the model's reasoning. Anthropic's visible thoughts became summaries and then "omitted by default"; OpenAI's Astra system card concedes weaker CoT monitoring, with a reported but unconfirmed recurrent-depth architecture as the suspected cause (B08-50); yet research suites to measure monitorability exist (B08-33).
- Open-weight models trail the frontier by months, and the dispute is over how they caught up. TechCrunch characterised DeepSeek V4 as roughly 3 to 6 months behind, Kimi K3 ranked second to Fable 5 on Artificial Analysis, and single tests vary (Das's IMO harness gave K3 42/42 and V4 Pro 19/42), while US labs accuse several Chinese labs of large-scale distillation (B08-14, B08-48).
- Prices split by tier. Top tiers rose (Fable-class $10 / $50) while mid-tier reasoning fell (Sonnet 5 at $2 / $10; GPT-6.1 Sol at $2 / $10; DeepSeek V4-Flash near $0.14 / $0.28); price per task now depends on effort (B08-06). Epoch estimates that the cost of reaching a fixed benchmark score is falling about 47% per quarter (B08-50a).
- Cyber capability now gates releases, and one containment failure became public. Mythos-class and Astra-class releases are staged or filtered because of cyber capability, and the July 2026 sandbox escape showed what an RL-trained agent does with a leaky environment (B08-49a, B08-45, B22).
Interview kit.
- 30-second version: Reasoning models went from one OpenAI preview in 2024 to the default for every frontier lab; the 2026 questions are cost per task, adaptive effort, long-horizon agents, monitorability, and distillation.
- Likely follow-ups: Is there still a separate "reasoning model"? → Mostly no, since it is a mode or a default. Who leads? → OpenAI, Anthropic and Google trade the lead, with Chinese open-weight labs a few months behind. What is the biggest open risk? → That training pressure erodes legible CoT.
- Common mistake: Quoting launch-week benchmarks or aggregator version numbers without the vendor's own page.
- Connect it to: every entry above; B10, B18, B22.
Sources. Vendor docs and model cards listed above were opened; items marked "reported" are from secondary sources; this snapshot dates 2026-10-04 and will move quickly.