Gemini 3 Pro scores 31.1% on ARC-AGI-2 and Deep Think 45.1%
Gemini 3 Pro launched with a 1M-token context and reported scores of Humanity's Last Exam 37.5% (no tools), GPQA Diamond 91.9%, SWE-bench Verified 76.2%, and LMArena 1501; Deep Think mode reported…
- Date
- 18 November 2025
- Who
- Google DeepMind
- Confidence
- High on launch claims; Medium on later numbers
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): Google DeepMind · Confidence: High on launch claims; Medium on later numbers Gemini 3 Pro launched with a 1M-token context and reported scores of Humanity's Last Exam 37.5% (no tools), GPQA Diamond 91.9%, SWE-bench Verified 76.2%, and LMArena 1501; Deep Think mode reported 41.0% on Humanity's Last Exam and 45.1% on ARC-AGI-2 with code execution, available first to safety testers and Ultra subscribers (Google). Gemini 3 Pro's ARC-AGI-2 score without Deep Think was 31.1%. The 2026 follow-ups took the ARC-AGI-2 gain further. Gemini 3.1 Pro (2026-02-19) was reported at a verified 77.1% on ARC-AGI-2, 44.4% on Humanity's Last Exam and 94.3% on GPQA Diamond, with three thinking_level settings at launch per VentureBeat (Google's docs now list four, minimal to high, with high the default on 3.1 Pro), where "high" works as a "mini Deep Think" (VentureBeat). A week earlier, on 2026-02-12, Google had upgraded Deep Think to a reported 84.6% on ARC-AGI-2, verified by the ARC Prize Foundation (B08-43b). ARC-AGI-2 (B08-21b), on which the released o3 scored under 3% at medium effort in April 2025 (B08-24), thus reached a verified 77.1% for a general model and 84.6% for Deep Think by February 2026. The successor benchmark is ARC-AGI-3, and B08-51 covers the current state. See also B23.