Google's Gemini 2.5 Pro thinking model takes first place on LMArena
Google's first frontier-scale reasoning model, announced as a "thinking model," took the top spot on LMArena (a human-preference leaderboard) by a significant margin and committed Google to building…
- Date
- 25 March 2025
- Who
- Google DeepMind
- Confidence
- High on claims; Medium on internals
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Landmark · Significance: 4/5 · Org(s): Google DeepMind · Confidence: High on claims; Medium on internals Primary sources: Gemini 2.5 announcement · Gemini 2.5 technical report, arXiv 2507.06261 · Gemini 2.5 Flash post
One-liner. Google's first frontier-scale reasoning model, announced as a "thinking model," took the top spot on LMArena (a human-preference leaderboard) by a significant margin and committed Google to building thinking into every model.
Why it happened. Google had the research (Brain's chain-of-thought and DeepMind's AlphaProof lineage, B07, B15) but shipped reasoning late. Flash Thinking came in December 2024, then nothing frontier-sized until this release 194 days after o1-preview and 64 after R1. I infer that what changed was applying RL at the scale of its largest sparse mixture-of-experts model, on top of the 1M-token multimodal Gemini base (technical report).
The idea. In the report's description, thinking models are trained with reinforcement learning to use extra compute at inference time, and a "Deep Think" variant adds parallel thinking, in which several hypotheses are explored at once before the model settles on an answer. A thinking budget makes the dial explicit (B08-06). Google says "we're building these thinking capabilities directly into all of our models" (announcement).
Results (company-reported). LMArena first place by a significant margin; Humanity's Last Exam (a hard expert-written question set) 18.8% without tools; SWE-bench Verified 63.8%; 1M-token context (2M promised).
How it spread. Gemini 2.5 became the template for the rest of Google's line, from 2.5 Flash with a thinking budget (+23 days) and 2.5 Deep Think (public 2025-08-01), which solved IMO problems (B08-34), to Gemini 3 (2025-11-18) and 3.x with thinking_level controls.
Why it mattered. By April 2025 all three US frontier labs and several Chinese labs had a flagship-class reasoning model, which ended the argument over whether reasoning was a boutique feature.
Nuance, controversy and myths. LMArena rank measures human preference and does not test reasoning. Google's earlier Flash Thinking was cheaper and earlier than R1 but did not set off a market reaction (B08-12).
Interview kit.
- 30-second version: Gemini 2.5 Pro is Google's RL-trained thinking model with a million-token context that topped the human-preference leaderboard in March 2025 and set Google's policy that every model thinks.
- Likely follow-ups: Why was Google late? → I infer that Flash Thinking was an experiment and the RL recipe had to be scaled to the Pro model. What is Deep Think? → Parallel thinking at higher compute; used for IMO gold.
- Common mistake: Saying Google "invented" reasoning models because of chain-of-thought; the 2022 CoT paper is a prompting method with no RL training (B07).
- Connect it to: B08-06, B08-34, B16.
Sources. Google pages opened; technical-report abstract opened (full text too large to fetch).