Google reports 84.6% on ARC-AGI-2 for the upgraded Gemini 3 Deep Think

Google introduced Gemini 3 Deep Think on 2025-12-04 for Google AI Ultra subscribers, reporting 41.0% on Humanity's Last Exam without tools and 45.1% on ARC-AGI-2 with code execution (Google), and…

Date
12 February 2026
Who
Google DeepMind
Confidence
Medium (company-reported; 9to5Google relays Google's post; the ARC leaderboard w
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): Google DeepMind · Confidence: Medium (company-reported; 9to5Google relays Google's post; the ARC leaderboard was not readable) Google introduced Gemini 3 Deep Think on 2025-12-04 for Google AI Ultra subscribers, reporting 41.0% on Humanity's Last Exam without tools and 45.1% on ARC-AGI-2 with code execution (Google), and upgraded it on 2026-02-12. Google reported 84.6% on ARC-AGI-2, "verified by the ARC Prize Foundation," 48.4% on Humanity's Last Exam without tools, a Codeforces Elo of 3455 and gold-medal-level IMO 2025 performance, available to Ultra subscribers and, by request, to enterprise API users (9to5Google).

A week later Gemini 3.1 Pro reached a verified 77.1% on the same benchmark (B08-42). So the best ARC-AGI-2 score went from 4% for the best listed system at launch (B08-21b) to 84.6% in under eleven months, at which point ARC-AGI-3 was weeks away (B08-44).

Google's docs list Gemini 3.5 to 3.8 Flash models (medium thinking by default) alongside gemini-3.1-pro-preview (high by default) (thinking docs); I found no primary source for a Gemini 3.5 Pro release (Backlog). Sources: Google, Deep Think · 9to5Google · Gemini thinking docs

Read it in the deep dive