Huawei and Xiaohongshu announce perfect IMO 2026 scores and Das's harness shows 42/42 runs

At the 67th IMO in Shanghai (papers 2026-07-15 and 07-16; 666 contestants, 7 human perfect scores) two systems, Huawei's "Celia" and Xiaohongshu's "dots-note-3.0," were announced as having scored…

Date
16 July 2026
Who
Huawei, Xiaohongshu (RedNote), Anthropic, OpenAI, Moonshot, Axiom Math
Confidence
Medium (AFP relays company claims; most per-model details come from one third-pa
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): Huawei, Xiaohongshu (RedNote), Anthropic, OpenAI, Moonshot, Axiom Math · Confidence: Medium (AFP relays company claims; most per-model details come from one third-party harness) At the 67th IMO in Shanghai (papers 2026-07-15 and 07-16; 666 contestants, 7 human perfect scores) two systems, Huawei's "Celia" and Xiaohongshu's "dots-note-3.0," were announced as having scored perfectly. Xiaohongshu said its solutions were submitted to IMO organisers for grading and that no large language model had previously scored perfectly under the IMO's official judging process; Huawei claimed 100%. AFP relays these company claims and quotes no IMO statement confirming the grading (AFP via France 24), so label them "company-announced, reportedly IMO-graded" and keep the 2025 IMO-president caveat in mind (B08-34).

Separately, investor Deedy Das ran models on the paper himself in a minimal agent harness (deedy/imo-2026). His repository shows Claude Fable 5 at 42/42 on a clean first pass (default high effort, 2.5 hours, $51.05); GPT-5.6 Sol at xhigh effort at 42/42 (3.8 hours, about $20.54) after reviewer-feedback repair rounds, from 39/42 on its first pass; and Kimi K3 at 42/42 (17.4 hours, about $31.40) after repair rounds, from 36/42. The same table gives GPT-5.6 Sol 28/42 at default effort and 30/42 at max effort (single-pass), Sol Pro 37/42, Meta's Muse Spark 1.1 26/42, DeepSeek V4 Pro 19/42 and Grok 4.5 13/42. It does not list Axiom Math's AxiomProver; that 42/42 comes from AFP's account of Das's statement. The graders are Claude-based agents (the repository says "strong but not authoritative"), not human medalists.

Compared with IMO 2025 (35/42, the gold line), olympiad-level problem solving went from gold-level to perfect in a year, at roughly $20 to $51 per run in Das's harness, though two of the three perfect runs needed repair rounds. The effort results (28, 39 and 30 out of 42 for default, xhigh and max effort) also show that more thinking does not always score higher (B08-06). The unofficial harness results should not be quoted as IMO results. See B16 and B23.

Read it in the deep dive