OpenAI and Google DeepMind reasoning models reach gold-medal level at IMO 2025

Two general-purpose reasoning models, working in natural language under contest time limits, scored 35 of 42 points at the 2025 International Mathematical Olympiad, the gold-medal line; both solved…

Date
19 July 2025
Who
OpenAI; Google DeepMind; Harmonic; ByteDance Seed
People
Alexander Wei, Noam Brown (OpenAI); the DeepMind IMO team
Confidence
High on scores; Medium on grading comparability
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Landmark · Significance: 5/5 · Org(s): OpenAI; Google DeepMind; Harmonic; ByteDance Seed · People: Alexander Wei, Noam Brown (OpenAI); the DeepMind IMO team · Confidence: High on scores; Medium on grading comparability Primary sources: DeepMind blog · Xena Project round-up (Kevin Buzzard) · Alexander Wei's announcement (X)

One-liner. Two general-purpose reasoning models, working in natural language under contest time limits, scored 35 of 42 points at the 2025 International Mathematical Olympiad, the gold-medal line; both solved five of six problems.

Why it happened. The IMO had been a grand-challenge target for years. In 2024, DeepMind's AlphaProof and AlphaGeometry 2 got silver-level (28 points) but needed formal proofs in Lean and up to three days of computation (B15, B16). The new results came from the reasoning stack of this chapter, meaning long chains of thought plus RL on proof-style tasks, with extra inference compute.

The idea and how it worked. OpenAI's "experimental reasoning LLM" was described as general purpose and not IMO-specific. It worked in two 4.5-hour sessions with no tools or internet, reading the official problems and writing natural-language proofs (reported). The proofs were graded by three former IMO medalists by consensus (reported) (LessWrong post reproducing OpenAI's announcement), and OpenAI said the model was experimental and would not be released for several months. Google's "advanced version of Gemini with Deep Think" (announced 2025-07-21 per Xena's round-up) used parallel thinking, RL on multi-step reasoning and theorem-proving data, and a curated corpus of high-quality solutions with IMO-specific guidance; it finished inside the 4.5-hour limit directly from natural-language statements, and IMO coordinators graded it (DeepMind).

Results. Both scored 35/42, with full marks on problems 1 to 5 and zero on problem 6, which no AI solved. ByteDance's formal Lean system Seed-Prover was first reported at silver level (2+7+7+7+7+0 after three days of computation) and then upgraded to a gold-level 5 of 6 (7+7+7+7+7+0) with additional attempts; its paper says it fully proved 5 of 6 problems (Xena, arXiv 2507.23726). Harmonic's Aristotle reached gold level in Lean, announced 2025-07-28 (Xena).

How it spread. Open-weight systems matched it within months. DeepSeekMath-V2 (arXiv 2025-11-27) claimed gold-level on IMO 2025 and 118/120 on Putnam 2024 using a trained verifier that rewards rigorous proofs as well as right answers (arXiv 2511.22570), and the high-compute DeepSeek-V3.2-Speciale reported 35/42 on IMO 2025 plus gold-level on IOI 2025 (492/600) and ICPC World Finals (10 of 12) (V3.2 paper, Table 4) (B08-43). In July 2026 the contest saw perfect scores (B08-48).

Why it mattered. It was the first widely accepted demonstration that general reasoning models can write human-readable olympiad proofs end to end, and the clearest evidence of the time-for-quality trade-off at scale (B16).

Nuance, controversy and myths. (1) Grading. The IMO President said the organisation could not validate methods, compute or human involvement and Buzzard criticised that companies "set and mark their own homework" (Xena). Google's grading was done by IMO coordinators, per Google. (2) Announcement etiquette. AI companies were asked to wait until 2025-07-28; OpenAI announced on 2025-07-19 after the closing ceremony and drew criticism from coordinators while Noam Brown said the only IMO contact had asked only for a post-ceremony announcement (reported, Zvi). (3) Compute and tooling. The amount of compute and any hidden scaffolding were not disclosed. (4) Olympiad problems are not research mathematics.

Interview kit.

  • 30-second version: In July 2025 OpenAI's and Google's reasoning models each solved five of six IMO problems in natural language within time limits, equalling the gold line; Google's was officially graded, OpenAI's was graded by former medalists.
  • Likely follow-ups: Why is this different from 2024? → No formal language, hours not days. Is it general? → Both say general-purpose reasoning models (Google gave Deep Think extra training data and IMO-specific general guidance). Did open models catch up? → DeepSeekMath-V2 and V3.2-Speciale within five months.
  • Common mistake: Saying the AI "won gold medals"; it was gold-medal-level, with no official medal.
  • Connect it to: B16, B15, B08-22.

Sources. DeepMind blog and Xena round-up opened; OpenAI-side details via secondary summaries (X and OpenAI pages not fetchable).

Read it in the deep dive