CAIS and Scale AI published Humanity's Last Exam, 2,500 expert questions behind many headline scores
Humanity's Last Exam (HLE) is a public set of 2,500 expert-written, closed-ended questions across dozens of subjects (mathematics, humanities, natural sciences), built by CAIS and Scale AI with…
- Date
- 24 January 2025
- Who
- Center for AI Safety (CAIS), Scale AI
- Confidence
- High
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): Center for AI Safety (CAIS), Scale AI · Confidence: High Humanity's Last Exam (HLE) is a public set of 2,500 expert-written, closed-ended questions across dozens of subjects (mathematics, humanities, natural sciences), built by CAIS and Scale AI with contributions from nearly 1,000 experts at over 500 institutions; the paper's arXiv v1 is dated 2025-01-24 (v11 on 2026-07-28), it was published in Nature on 2026-01-28, and a private held-out set is kept to detect overfitting (arXiv 2501.14249, lastexam.ai). The project began as a call for questions on 2024-09-16 with a $500,000 prize pool (Scale AI). It earns an entry because it supplies many of this chapter's headline numbers (R1-0528 17.7%, Gemini 2.5 Pro 18.8%, Gemini 3 Pro 37.5%, Grok 4 Heavy 50.7%, Gemini 3 Deep Think 48.4%) and because those numbers are not comparable unless the condition is stated, whether with or without tools, full set or text-only subset, single model or parallel agents. A hard benchmark saturates quickly under RL-trained reasoning (my inference), and the maintainers themselves wrote that models might exceed 50% by the end of 2025. Related reading is B23. Sources: arXiv · lastexam.ai · Scale AI