METR's time-horizon paper measures AI by the length of tasks it completes
METR proposed measuring AI progress by the length of human task, in expert-human time, that a model completes with 50% success, found this horizon doubling about every seven months since 2019, and…
- Date
- 18 March 2025
- Who
- METR
- People
- Thomas Kwa, Ben West, Joel Becker and colleagues (26 listed authors in v4)
- Confidence
- High on the paper and METR's notes; Medium on 2026 model numbers (the page I rea
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Landmark · Significance: 4/5 · Org(s): METR · People: Thomas Kwa, Ben West, Joel Becker and colleagues (26 listed authors in v4) · Confidence: High on the paper and METR's notes; Medium on 2026 model numbers (the page I read lists few models) Primary sources: arXiv 2503.14499 (v1 2025-03-18, v4 2026-07-10; NeurIPS 2025) · METR time-horizons page (updated 2026-05-08) · METR note, 2026-01-22
One-liner. METR proposed measuring AI progress by the length of human task, in expert-human time, that a model completes with 50% success, found this horizon doubling about every seven months since 2019, and gave reasoning-era agents their standard yardstick.
Why it happened. Benchmarks such as AIME and GPQA saturate within months of a reasoning model's release (B08-11a, B08-48) and say little about real work. METR, an evaluation nonprofit, wanted a unit with human meaning, the time a task takes a person. The paper appeared 187 days after o1-preview. I infer that reasoning models were making multi-step tasks newly feasible, which made a measure built on task length informative.
The idea. Time human experts on a task set, run models on the same tasks, and report the human time at which a model's success rate falls to 50%. How it works. The authors timed humans with relevant expertise on RE-Bench, HCAST and 66 novel shorter tasks, and read each model's 50% horizon from its success rate against human time.
Results. In the paper, frontier models such as Claude 3.7 Sonnet had a 50% horizon of around 50 minutes. The horizon had doubled roughly every seven months since 2019, perhaps faster in 2024, and the authors attribute growth mainly to greater reliability and ability to adapt to mistakes, with better logical reasoning and tool use. If results generalise to real-world software, they extrapolate that within five years AI could automate many software tasks that take humans a month. A METR note of 2026-01-22 puts Claude Opus 4.5 at about 4 hours 49 minutes (95% confidence interval 1h49m to 20h25m) and the long-run doubling time at 6 to 7 months. METR's time-horizons page, updated 2026-05-08, lists Claude Opus 4.6, GPT-5.4, Gemini 3.1 Pro and Claude Mythos Preview (early), and says measurements above 16 hours are unreliable with the current task suite; as fetched, it did not list Opus 4.7, GPT-5.5, GPT-5.6, Fable 5 or GPT-6, so later models' horizons are not established here (see the Backlog).
How it spread. The page carries measurements for models from OpenAI, Anthropic and Google, which makes it a cross-lab yardstick, and the 50% horizon is the quantity that B08-51 uses for long-horizon progress.
Why it mattered. It turned an argument about agent capability into a trend line, the nearest thing to a quantitative payoff statement for reasoning plus tool use, and it links this chapter to B10 and B24.
Nuance, controversy and myths. METR cautions in its note of 2026-01-22 that the horizon does not measure "the length of time AIs can work independently". It measures the amount of serial human labour a model can replace at 50% success. Error bars span roughly a factor of two in each direction; horizons differ by orders of magnitude across domains (visual computer-use tasks 40 to 100 times lower than software and research tasks); the tasks omit real-world context and collaboration; and baseliner choices could shift measurements by more than 25%. A 50% horizon says nothing about the length at which a model is reliable. The paper's title changed from "Measuring AI Ability to Complete Long Tasks" (v1) to "...Long Software Tasks" (v4).
Interview kit.
- 30-second version: METR times human experts on tasks and finds the task length at which each AI model succeeds half the time; that horizon was about 50 minutes for Claude 3.7 Sonnet in early 2025 and has doubled roughly every six to seven months.
- Likely follow-ups: Does a 4h49m horizon mean Opus 4.5 can work autonomously for five hours? → No, it means tasks that take a human expert that long are completed about half the time, with a confidence interval from 1h49m to 20h25m. Is progress accelerating? → METR reports a 6 to 7 month long-run doubling, possibly faster since 2024. Why is it a reasoning-era metric? → The paper credits reliability, error correction, logical reasoning and tool use, the traits RL-trained reasoning models improved.
- Common mistake: Reading the horizon as autonomous run time or as a 95%-reliability figure.
- Connect it to: B10, B23, B24, B08-51.
Sources. [1] arXiv abstract and version history, opened. [2] METR time-horizons page, opened (numeric values were not in the text returned). [3] METR note, opened.