ARC Prize launches ARC-AGI-3, where humans scored 100% and the best AI system 0.51%

After ARC-AGI-2 reached a verified 77.1% for a general model (Gemini 3.1 Pro, 2026-02-19) and a reported 84.6% for Gemini 3 Deep Think (B08-43b), the ARC Prize Foundation launched ARC-AGI-3 on…

Date
25 March 2026
Who
ARC Prize Foundation
Confidence
High (ARC Prize pages opened)
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): ARC Prize Foundation · Confidence: High (ARC Prize pages opened) After ARC-AGI-2 reached a verified 77.1% for a general model (Gemini 3.1 Pro, 2026-02-19) and a reported 84.6% for Gemini 3 Deep Think (B08-43b), the ARC Prize Foundation launched ARC-AGI-3 on 2026-03-25. It has hundreds of hand-built, interactive, turn-based game environments with no instructions, where agents must explore, infer the rules and win. Humans scored 100% and the best frontier AI system scored 0.51% at launch, with over $2 million in prizes (ARC Prize announcement).

That page lists no per-model scores, and per-model launch figures seen in secondary coverage could not be found on the ARC or DataCamp pages, so none are quoted here. By 2026-09-03 OpenAI said GPT-6 Astra "saturates" it with 99.9%. ARC Prize's own post of that date reports 62.7% on the semi-private set under its standard, provider-neutral harness (about $26,098) and 99.9% under a provider-adapter harness (about $18,817, at high reasoning effort). The standard harness lets the model carry forward notes it chooses to keep; the adapter preserves opaque reasoning state between requests and compacts long conversations, and ARC reports it ran about 3.66x faster and used 49% fewer tokens (ARC Prize, Astra).

Harness design now matters as much as the model, and the gap between the two numbers is an interviewer trap. For lineage see B08-08, B23.

Read it in the deep dive