Epoch AI finds the price of reaching a fixed benchmark score fell about 47% per quarter
Epoch AI estimates that the cost of reaching a given benchmark score has fallen about 47% per quarter (roughly 13x a year) since 2023.
- Date
- 22 September 2026
- Who
- Epoch AI
- Confidence
- Medium (Epoch's analysis, read as a summary of the page; underlying data not ins
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): Epoch AI · Confidence: Medium (Epoch's analysis, read as a summary of the page; underlying data not inspected) Epoch AI estimates that the cost of reaching a given benchmark score has fallen about 47% per quarter (roughly 13x a year) since 2023. It measures the price of the cheapest model on the empirical Pareto frontier that reaches a fixed accuracy threshold, with raw per-token prices excluded, across AIME (OTIS Mock), GPQA Diamond, FrontierMath Tiers 1 to 3, chess puzzles and mystery-game puzzles. For example, for GPQA Diamond at 75%, o3 (January 2025) cost about $0.30 per question and GPT-5.6 Luna (July 2026) about $0.0004, a 725-fold cut in 18 months.
Maths benchmarks fell 50 to 52% a quarter and game puzzles 39 to 43%; declines are fastest right after a score first becomes state of the art (66% a quarter, easing to 32% two years later), so top performance carries a short premium. Epoch gives four caveats. Models may be tuned to known tests, there are only three years of data, users do not always switch to the cheapest model, and estimates range from 42.9% to 58% across statistical methods (Epoch). It matters here because it quantifies the price split described in B08-51 and the thesis of B08-06 that thinking is sold as a price ladder, since the frontier gets pricier while a fixed capability gets much cheaper. See also B18. Sources: Epoch