OpenAI previewed o3, which scored 87.5% on ARC-AGI-1 at about $4,560 a task

On the last of OpenAI's "12 days" streams, o3 was previewed with 87.5% on ARC-AGI-1 and 25.2% on FrontierMath, reached by spending far more inference compute than o1.

Date
20 December 2024
Who
OpenAI; ARC Prize Foundation; Epoch AI
People
Sam Altman, Mark Chen, Hongyu Ren, François Chollet
Confidence
High on what was announced; Medium-Low on what the announced numbers mean for th
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Landmark · Significance: 5/5 · Org(s): OpenAI; ARC Prize Foundation; Epoch AI · People: Sam Altman, Mark Chen, Hongyu Ren, François Chollet · Confidence: High on what was announced; Medium-Low on what the announced numbers mean for the released model Primary sources: ARC Prize on the o3 breakthrough on ARC-AGI-Pub · ARC Prize analysis of the released o3 and o4-mini · Epoch on OpenAI and FrontierMath · Deliberative alignment, arXiv 2412.16339

One-liner. On the last of OpenAI's "12 days" streams, o3 was previewed with 87.5% on ARC-AGI-1 and 25.2% on FrontierMath, reached by spending far more inference compute than o1.

Why it happened. o1-preview was about 100 days old. OpenAI wanted to show that the RL-on-CoT axis kept scaling, and that it worked beyond AIME-style problems. I infer that the timing also answered a crowded news cycle, since Google's Flash Thinking had appeared the day before. OpenAI skipped the name "o2" because of a trademark conflict with the telecom brand O2 (reported; The Decoder).

The idea. I infer that the recipe is the same as o1, with more RL and many more tokens (and, for the headline ARC runs, many parallel samples) at inference. OpenAI released no architecture or training details; what is documented is that OpenAI described o3 as a 10x scale-up in training compute over o1 (Epoch reads this as RL compute, see B08-24), that the two ARC settings used 6 and 1,024 samples, and that OpenAI told ARC the tested o3 had been trained on 75% of the ARC-AGI-1 public training set (ARC Prize). It did publish a companion paper, "Deliberative Alignment" (arXiv v1 2024-12-20), in which the model is taught a written safety specification and reasons over it in its chain of thought before answering (paper).

Results (company-reported unless noted). On FrontierMath (Epoch AI's research-level problem set), o3 scored 25.2% versus about 2% for previous systems. Codeforces rating about 2727. SWE-bench Verified (real GitHub software-engineering tasks) 71.7% versus 48.9% for o1. AIME 2024 96.7%. GPQA Diamond 87.7% (coverage summarising the livestream). The ARC Prize Foundation tested "o3-preview" itself on ARC-AGI-1 (the Abstraction and Reasoning Corpus, a set of novel grid puzzles that test adaptation to unseen tasks; results post) and reported the results below.

ARC-AGI-1, semi-private set (100 tasks)ScoreSamplesTokensCost per task
"High-efficiency" (within the $10k prize budget)75.7%633.5Mabout $26
"Low-efficiency" (172x compute)87.5%1,0245.7Babout $4,560

For context, GPT-3 scored 0% on ARC-AGI in 2020 and GPT-4o 5% in 2024. ARC called the result a genuine step change in task adaptation, and François Chollet said plainly that o3 was not yet AGI, because easy tasks still failed and a harder successor benchmark, ARC-AGI-2, was coming.

How it spread. Open labs had no access to o3's weights and still copied the framing quickly, since reasoning effort and test-time compute were now a headline metric. R1 reported results comparable to OpenAI's o1 (not o3), ahead on some benchmarks and behind on others, 31 days after this announcement; Anthropic's Claude 3.7 and Google's Gemini 2.5 Flash added token-budget controls (B08-06); ARC's leaderboard now plots score against cost per task (B23).

Why it mattered. It converted "thinking longer" into a dollar-per-task curve. It also showed the ceiling of that approach, since the 87.5% run cost about 172 times the 75.7% run for 11.8 more points.

Nuance, controversy and myths.

  • The model announced is not the model shipped. ARC states the production o3 (released 2025-04-16) "uses a different model from the o3-preview," is multimodal, gets far less test-time compute, was tuned for chat and product use, and was not trained directly on ARC-AGI-1 data as the preview reportedly was. Released o3 scored 41% (low) and 53% (medium) on ARC-AGI-1 at about $1.22 and $2.52 a task, and under 3% on ARC-AGI-2 at medium effort (ARC Prize, 2025-04-22). The high-effort runs there were incomplete.
  • FrontierMath conflict. OpenAI funded and owned the 300 problems commissioned for FrontierMath and had their statements and solutions. Epoch said it was still finalising a 50-problem holdout whose solutions OpenAI would not receive, so no clean held-out test existed for the December run; Epoch got OpenAI's permission to disclose the arrangement ahead of the o3 announcement but made it public only around that time, and later said it had not communicated the relationship clearly enough (Epoch; TechCrunch, 2025-01-19). When Epoch ran the released o3 in April 2025 it measured roughly 10% on an updated, 290-question version, versus the 25% announced under aggressive compute settings (reported; see coverage, primary Epoch post not retrieved).
  • "o3 solved ARC-AGI" overstates it, because it was the public/semi-private ARC-AGI-1 set, at a cost no user could pay, with ARC explicitly disclaiming AGI.

Interview kit.

  • 30-second version: o3 is OpenAI's second-generation reasoning model. At launch it posted 87.5% on ARC-AGI-1 by spending thousands of dollars of compute per task, which showed that reasoning accuracy tracks inference spend, and the model that shipped months later was a cheaper, different one.
  • Likely follow-ups: Was it AGI? → ARC's authors said no. Why did the released o3 score lower? → Different model, far less compute, plus benchmark-specific preview tuning. Why does the FrontierMath result carry an asterisk? → OpenAI funded it and had access to most problems.
  • Common mistake: Quoting "$20 per task" for the 75.7% run (ARC's table says about $26) or quoting the 87.5% result as the released model's.
  • Connect it to: B08-01, B23, B16, B18.

Sources. ARC Prize posts above (opened); Epoch statement (opened); Deliberative alignment (opened); livestream numbers via secondary coverage because OpenAI's pages return HTTP 403 to the fetcher.

Read it in the deep dive