OpenAI launches GPT-5 with a router that decides when to think

OpenAI folded its reasoning (o-series) and chat (GPT-4o) lines into one product made of a fast model, a deeper "thinking" model and a real-time router that picks between them.

Date
7 August 2025
Who
OpenAI
Confidence
High on design; Medium on user-impact claims
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Landmark · Significance: 4/5 · Org(s): OpenAI · Confidence: High on design; Medium on user-impact claims Primary sources: GPT-5 System Card, arXiv 2601.03267 (v1 2025-12-19; originally published 2025-08-07 on OpenAI's site) · Simon Willison's notes · Fortune on the router backlash

One-liner. OpenAI folded its reasoning (o-series) and chat (GPT-4o) lines into one product made of a fast model, a deeper "thinking" model and a real-time router that picks between them.

Why it happened. The intent was on the record months earlier, because in the 2025-02-12 roadmap post Altman called GPT-4.5 the last non-chain-of-thought model and said GPT-5 would integrate o3, which would no longer ship separately (B08-17c). I infer that the practical motives were a confusing ChatGPT menu (GPT-4o, o3, o4-mini...) and expensive reasoning compute, though OpenAI has not itemised them. Anthropic had shown a single hybrid model (Claude 3.7) and Google a thinking default (Gemini 2.5). OpenAI's answer was to hide the choice in a system. It had a fast model for most questions, a deeper reasoning model for harder ones, and a real-time router that chooses using conversation type, complexity, tool needs and explicit user intent (system card abstract).

The idea. The API offers gpt-5, gpt-5-mini and gpt-5-nano, each with four reasoning levels including the new "minimal". Reasoning tokens are generated by default unless minimal is selected, context is up to 272K input and 128K output tokens, and gpt-5 costs $1.25 / $10 per million input / output tokens, half GPT-4o's input price (Willison). ChatGPT has Auto, Fast, Thinking and Thinking-mini modes, with Thinking limited to 3,000 messages a week after the first update (reported). The system card classified gpt-5-thinking as high-capability in biology and chemistry and describes "safe-completions" training.

Results. OpenAI claimed lower hallucination and sycophancy than predecessors. Simon Willison reported that after two weeks of testing the model "rarely screws up" and generally felt competent, and flagged a 56.8% prompt-injection attack success rate in OpenAI's own tests.

How it spread. The router idea was echoed in DeepSeek's hybrid V3.1 (B08-37) and Anthropic's adaptive thinking (2026-02, B08-06); OpenAI iterated through GPT-5.1 to GPT-5.6 and then GPT-6 (B08-51).

Why it mattered. It moved reasoning from a model choice to a system behaviour and made "how long to think" a routing problem, with latency and cost consequences (B18).

Nuance, controversy and myths. The launch was bumpy. The router malfunctioned for part of launch day, making GPT-5 look weaker, and users objected to losing model choice, so OpenAI restored GPT-4o access, fixed routing and raised limits (Fortune). "GPT-5 is a single model" is wrong, because it is a system. "GPT-5 is not a reasoning model" is also wrong, because its thinking variant continues the o-series (inference from the system card's description).

Interview kit.

  • 30-second version: GPT-5 is a router over a fast model and a thinking model, with effort levels from minimal up; it unified OpenAI's lineup and made thinking depth an automatic decision.
  • Likely follow-ups: Where did the o-series go? → Into gpt-5-thinking (inference from the system card). Why did the launch backfire? → Router outage and loss of model choice.
  • Common mistake: Calling it "one model with a switch."
  • Connect it to: B08-06, B08-19, B05.

Sources. System-card abstract, Willison and Fortune opened; OpenAI's launch page not fetchable (403).

Read it in the deep dive