ARC-AGI-2 launch
The ARC Prize Foundation launched ARC-AGI-2 on 2025-03-24, keeping the "easy for humans, hard for AI" principle with harder, more compositional tasks and adding efficiency (cost per task) as a…
- Date
- 24 March 2025
- Who
- ARC Prize Foundation
- Confidence
- High
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): ARC Prize Foundation · Confidence: High The ARC Prize Foundation launched ARC-AGI-2 on 2025-03-24, keeping the "easy for humans, hard for AI" principle with harder, more compositional tasks and adding efficiency (cost per task) as a reported measure. Every task had been solved by at least two humans in two attempts or fewer, and a human panel averaged 60%. At launch (pass@2) pure LLMs scored 0%, o1-pro about 1%, DeepSeek R1 0.3%, o3-mini-high 0% and the o3-preview low-compute setting 4% (ARC Prize). ARC Prize 2025 offered $1 million, including a $700,000 grand prize for a solution above 85%, on Kaggle from 2025-03-26 to 2025-11-03. It belongs here because the chapter cites ARC-AGI-2 repeatedly, with the released o3 under 3% (2025-04), Gemini 3 Pro at 31.1% and Deep Think at 45.1% (2025-11), Gemini 3.1 Pro at a verified 77.1% and Google's upgraded Deep Think at 84.6% (2026-02, B08-43b). I infer that the climb, from near zero to the mid-80s in under a year for the best systems, left little headroom, and it is the context for ARC-AGI-3. Sources: ARC Prize announcement