Claude 3.7 Sonnet launches as a hybrid reasoning model with visible thinking
Anthropic's first reasoning model put thinking and non-thinking behaviour in one model with a token-budget dial, showed the raw thoughts, and was tuned for real-world coding, with less emphasis on…
- Date
- 24 February 2025
- Who
- Anthropic
- Confidence
- High
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Landmark · Significance: 4/5 · Org(s): Anthropic · Confidence: High Primary sources: Claude 3.7 Sonnet and Claude Code · Visible extended thinking · The Batch summary
One-liner. Anthropic's first reasoning model put thinking and non-thinking behaviour in one model with a token-budget dial, showed the raw thoughts, and was tuned for real-world coding, with less emphasis on contest maths.
Why it happened. Anthropic was the last of the big three US labs without a reasoning product at the time of R1. I infer that its design was a product-philosophy choice made independently of R1, since 35 days is too short to respond. Anthropic wanted one model that can answer instantly or think at length, and called it the first hybrid reasoning model "on the market" (announcement). The announcement also said Anthropic had shifted emphasis from competition maths and computer-science problems to tasks that reflect how businesses use LLMs, the bet that fed B10, and Claude Code launched in the same post as a research preview.
The idea. Users or developers toggle "extended thinking" and, via the API, set a thinking-token budget of up to 128K output tokens; thinking tokens are billed at the output price ($15 per million) with unchanged $3 per million input (The Batch). Anthropic's companion post lists the reasons to show the thoughts (trust, alignment research, interest to users) and the costs (the thoughts may be unfaithful to the actual computation, they lack the model's character, and adversaries could mine them), and shows accuracy on maths rising with thinking tokens and, at 256 parallel samples with a learned scorer, GPQA Diamond 84.8% (visible-thinking post).
Results (company-reported). SWE-bench Verified 63.7% with basic bash and file-editing tools and 70.3% with a high-compute scaffold, both on the 489 of 500 tasks that run on Anthropic's infrastructure, so they are not directly comparable to full-500 scores (per the announcement page); AIME 2024 80.0% and GPQA Diamond 84.8% in parallel extended-thinking mode with a 64K budget, versus o3-mini's 87.3% on AIME (The Batch).
How it spread. Others copied or re-derived the hybrid framing over the following months, with Google's 2.5 Flash (2025-04-17, 52 days later), Qwen3 (+64), DeepSeek V3.1's hybrid and OpenAI's GPT-5 router (+164, B08-36). Anthropic's own follow-through reached its end state in 2026, with adaptive thinking by default and always-on thinking in Fable-class models (B08-06).
Why it mattered. It treated reasoning as a capability inside a general model, where earlier reasoning models were separate model families, and aimed it at agentic coding. The same post introduced Claude Code, Anthropic's agentic coding tool, as a research preview (B10). It also raised the best-studied open question on visible thoughts. Anthropic's own faithfulness study of this model (B08-23) found that it mentioned hints it had used only 25% of the time.
Nuance, controversy and myths. Calling it "the first hybrid" is a company claim. OpenAI's o1 API had effort levels earlier (2024-12-17, B08-06), though it had no unified model with an explicit token budget. Google's Flash Thinking was a separate experimental model, and Google's own "fully hybrid" model, Gemini 2.5 Flash, came 52 days after Claude 3.7. Raw thoughts were shown in this release but later became summarised (Claude 4) and, from April 2026 (Mythos Preview, Opus 4.7), omitted by default.
Interview kit.
- 30-second version: Claude 3.7 Sonnet is one model with two modes (fast and extended thinking), exposes a thinking-token budget, shows the thoughts, and was tuned for software engineering, with less emphasis on contest maths.
- Likely follow-ups: Why a hybrid? → One model to serve and one product surface; thinking is a capability. Why show thoughts? → Trust, alignment research, and research interest, with caveats on faithfulness. How is thinking billed? → As output tokens.
- Common mistake: Saying Anthropic had no reasoning research before 3.7; the company's own post describes serial and parallel test-time scaling curves.
- Connect it to: B08-06, B08-23, B10, B06.
Sources. Anthropic pages opened; The Batch summary opened for budget, pricing and competitor numbers.