CAIS and Scale AI published Humanity's Last Exam, 2,500 expert questions behind many headline scores

Humanity's Last Exam (HLE) is a public set of 2,500 expert-written, closed-ended questions across dozens of subjects (mathematics, humanities, natural sciences), built by CAIS and Scale AI with…

Date
24 January 2025
Who
Center for AI Safety (CAIS), Scale AI
Confidence
High
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): Center for AI Safety (CAIS), Scale AI · Confidence: High Humanity's Last Exam (HLE) is a public set of 2,500 expert-written, closed-ended questions across dozens of subjects (mathematics, humanities, natural sciences), built by CAIS and Scale AI with contributions from nearly 1,000 experts at over 500 institutions; the paper's arXiv v1 is dated 2025-01-24 (v11 on 2026-07-28), it was published in Nature on 2026-01-28, and a private held-out set is kept to detect overfitting (arXiv 2501.14249, lastexam.ai). The project began as a call for questions on 2024-09-16 with a $500,000 prize pool (Scale AI). It earns an entry because it supplies many of this chapter's headline numbers (R1-0528 17.7%, Gemini 2.5 Pro 18.8%, Gemini 3 Pro 37.5%, Grok 4 Heavy 50.7%, Gemini 3 Deep Think 48.4%) and because those numbers are not comparable unless the condition is stated, whether with or without tools, full set or text-only subset, single model or parallel agents. A hard benchmark saturates quickly under RL-trained reasoning (my inference), and the maintainers themselves wrote that models might exceed 50% by the end of 2025. Related reading is B23. Sources: arXiv · lastexam.ai · Scale AI

Read it in the deep dive