OpenAI, Anthropic, Google and others made thinking time a setting, from effort levels to adaptive thinking

Within a year every lab turned "how long the model thinks" into an API setting, then, from 2026, took that control away from developers and let the model choose.

Date
17 December 2024
Who
OpenAI, Anthropic, Google, Alibaba, DeepSeek, Moonshot
Confidence
High for API facts opened in vendor docs; Medium for pricing history
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Landmark · Significance: 4/5 · Org(s): OpenAI, Anthropic, Google, Alibaba, DeepSeek, Moonshot · Confidence: High for API facts opened in vendor docs; Medium for pricing history Primary sources: Anthropic extended-thinking docs · Claude Platform release notes · Gemini thinking docs · OpenAI reasoning guide · Gemini 2.5 Flash post, 2025-04-17

One-liner. Within a year every lab turned "how long the model thinks" into an API setting, then, from 2026, took that control away from developers and let the model choose.

Why it happened. Reasoning tokens cost real money and real latency, and the right amount varies by prompt by two or more orders of magnitude (a greeting needs none; an olympiad proof may take hundreds of thousands). Labs needed a way to sell the same model to latency-sensitive chat users and to batch-research users, and to avoid "overthinking," a failure mode in which models burn tokens on easy questions (noted by DeepSeek itself in its limitations).

The idea and how it evolved.

DateControlLabWhat it did
2024-12-17reasoning_effort low / medium / highOpenAI (o1 API)First provider-level thinking dial, shipped with the o1-2024-12-17 snapshot; reasoning tokens billed as output (Neowin and an OpenAI developer-forum thread, both known from search extracts only)
2025-01-31o3-mini low / medium / high; free-tier accessOpenAIEffort became a ChatGPT menu item too (B08-16)
2025-02-24budget_tokens (up to 128K)AnthropicDeveloper sets a token ceiling; thinking billed at output-token price (B08-19)
2025-04-17thinking_budget 0-24,576 tokens, thinking on/offGoogle (2.5 Flash)"Fully hybrid"; the model is trained to decide how long to think; launch pricing split thinking vs non-thinking output ($3.50 vs $0.60 per million), later unified at $2.50 on general availability (discussion of the split)
2025-04-29/think and /no_think, thinking budgetAlibaba (Qwen3)Open-weight hybrid model; reversed in July (B08-20)
2025-08-07minimal effort plus a router in ChatGPTOpenAI (GPT-5)Four levels; thinking on by default unless minimal (Willison); see B08-36
2025-11-24effort parameterAnthropic (Opus 4.5)At medium effort Opus 4.5 reportedly matched Sonnet 4.5's best SWE-bench score with 76% fewer output tokens (coverage)
2026-02-05Adaptive thinking; budget_tokens deprecatedAnthropic (Opus 4.6)The model decides whether and how much to think; effort replaces budgets (release notes)
2026-04-07 / 04-16Thinking display default "omitted" (Mythos Preview, then Opus 4.7); xhigh effort (between high and max) and manual budgets rejected with HTTP 400 on Opus 4.7AnthropicOmitted blocks carry only an encrypted signature; xhigh is aimed at agentic and coding sessions over 30 minutes (docs, release notes)
2026-05-28Adaptive thinking triggers reasoning only when a turn needs itAnthropic (Opus 4.8)Fewer wasted thinking tokens at the same effort than Opus 4.7 (release notes)
2026-06-09Always-on adaptive thinking (Fable 5, Mythos 5); the "omitted" display default carries overAnthropicThinking cannot be disabled; readable summaries are opt-in via display: "summarized"; the raw chain of thought is never returned (release notes)
2026-06-30 / 07-24Adaptive thinking on by default and manual budgets removed (Sonnet 5); thinking on by default (Opus 5)AnthropicSonnet 5 at $2 / $10 per million tokens (introductory, standard from 2026-08-10); Opus 5 at $5 / $25 (release notes)
2026-08-18 (beta header)display: "updates" returns short progress notes between tool calls instead of thoughtsAnthropicReasoning blocks stay empty; the notes come back as readable text (docs)
2026 (docs as of 2026-10)thinking_level minimal, low, medium, high; none to max effort ladderGoogle (Gemini 3.x); OpenAI (GPT-5.x/6.x)Gemini's docs list four levels with medium the default on the 3.5-3.8 Flash models and high on gemini-3.1-pro-preview; OpenAI documents seven effort values and some models reject none (Gemini docs, OpenAI docs)
2026-07Reasoning always on, one max level at launchMoonshot (Kimi K3)Reasoning as the only mode (Willison)

Pricing implications. Hidden thinking tokens are billed as output tokens at every major vendor, so the list price per million tokens understates the cost per task for reasoning models; OpenAI itself advises reserving at least 25,000 tokens for reasoning when starting out (OpenAI docs). The tier structure moved from o1 pro mode at $200 a month (B08-05) to o3-pro at $20 / $80 per million input / output tokens and a simultaneous 80% cut of o3 to $2 / $8 on 2025-06-10 (o3 cut, o3-pro price); to GPT-5 at $1.25 / $10 per million (2025-08-07, Willison); and, at the top of 2026, Anthropic's Fable-class models at $10 / $50 against $2 / $10 for Sonnet 5 (release notes, reported pricing). I infer that the quality ladder is now sold as a price ladder, with the effort setting acting as a second, finer one.

How it spread. Each lab copied the others within one release cycle. OpenAI's effort levels (2024-12) preceded Anthropic's budget (+69 days), Google's budget (+121 days) and the Qwen toggle (+133 days); Anthropic's adaptive mode (2026-02) was followed by Gemini's dynamic default and OpenAI's effort ladder. The open-weight labs went in different directions. Qwen3 moved from hybrid toggles to separate Instruct and Thinking models within three months, and DeepSeek's V4 merged its V and R lines into one model (B08-47).

Why it mattered. It turned test-time compute into something customers buy by the unit. The quality-versus-cost trade-off that o1's blog showed as a curve is now a setting on every request, and the per-task cost of agents is dominated by it (B10, B18).

Nuance, controversy and myths. Anthropic's docs treat a "thinking budget" as a target and say Claude may stop reasoning well before the budget is used, with max_tokens as the hard ceiling (the budget can exceed max_tokens only with interleaved thinking). "More thinking is always better" is false, since a 14-author team documented tasks where longer reasoning lowers accuracy (Inverse Scaling in Test-Time Compute, arXiv 2507.14417, 2025-07-19) and Apple's puzzle study found effort falling at high difficulty (B08-28). Display policy diverged from the 2025 norm of showing thoughts. Anthropic's visible thinking at launch (B08-19) became summaries in Claude 4 and, from Mythos Preview and Opus 4.7 in April 2026, "omitted by default"; Anthropic's docs state that no display setting returns the raw chain of thought. The same docs say the encrypted signature carries the full thinking for multi-turn continuity, a design that B08-49b later attacked.

Interview kit.

  • 30-second version: A reasoning model's quality is a function of how many hidden tokens it spends, so vendors expose that as effort levels or token budgets, bill the tokens as output, and are now moving to adaptive modes where the model picks.
  • Likely follow-ups: Is a thinking token the same price as an output token? → Yes at OpenAI, Anthropic and Google (Google briefly split the price). Why drop budgets? → Anthropic's docs note that changing a budget invalidates prompt-cache breakpoints, and adaptive modes let the model decide (my inference is that removing a tuning burden is a main motive). Why hide thoughts again? → Distillation risk and safety monitoring; Anthropic's Fable 5 even added a "reasoning_extraction" refusal category for requests that try to duplicate its outputs (release notes, 2026-06-09; B08-14).
  • Common mistake: Comparing list prices per million tokens across a reasoning and non-reasoning model without counting reasoning tokens.
  • Connect it to: B08-01, B08-36, B18.

Sources. Vendor documentation pages above were opened and read (Anthropic, Google, OpenAI); the Claude release-notes table is the primary record for all 2026 Anthropic dates; coverage links for 2025 pricing are secondary.

Read it in the deep dive