ARC-AGI-2 launch

The ARC Prize Foundation launched ARC-AGI-2 on 2025-03-24, keeping the "easy for humans, hard for AI" principle with harder, more compositional tasks and adding efficiency (cost per task) as a…

Date
24 March 2025
Who
ARC Prize Foundation
Confidence
High
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): ARC Prize Foundation · Confidence: High The ARC Prize Foundation launched ARC-AGI-2 on 2025-03-24, keeping the "easy for humans, hard for AI" principle with harder, more compositional tasks and adding efficiency (cost per task) as a reported measure. Every task had been solved by at least two humans in two attempts or fewer, and a human panel averaged 60%. At launch (pass@2) pure LLMs scored 0%, o1-pro about 1%, DeepSeek R1 0.3%, o3-mini-high 0% and the o3-preview low-compute setting 4% (ARC Prize). ARC Prize 2025 offered $1 million, including a $700,000 grand prize for a solution above 85%, on Kaggle from 2025-03-26 to 2025-11-03. It belongs here because the chapter cites ARC-AGI-2 repeatedly, with the released o3 under 3% (2025-04), Gemini 3 Pro at 31.1% and Deep Think at 45.1% (2025-11), Gemini 3.1 Pro at a verified 77.1% and Google's upgraded Deep Think at 84.6% (2026-02, B08-43b). I infer that the climb, from near zero to the mid-80s in under a year for the best systems, left little headroom, and it is the context for ARC-AGI-3. Sources: ARC Prize announcement

Read it in the deep dive