Long-horizon autonomous agents
Agents that work autonomously for hours to days, often in parallel teams, tracked by the length of task they complete rather than benchmark accuracy.
Progress is best tracked by task length, and METR's 50%-success time horizon has doubled roughly every 3-4 months since 2024. RL on long tasks plus harnesses (agent teams, swarms, context compaction) extend it, and the camp expects autonomous multi-day work to become routine.
Where it stands. Task horizons are growing fast but poorly measured, and METR's own top-end estimates swing from 11 to 270+ hours. Parallel agent teams are in use; independent evidence of reliable multi-day autonomy remains thin.
Evidence
For
- METR Time Horizon 1.1 (2026-01-29): doubling about every 131 days since 2023 and 89 days since 2024; Claude Opus 4.5 50% horizon ~320 minutes versus GPT-5 at 214 and o3 at 121.
- Sixteen parallel Claude Opus 4.6 agents wrote a 100,000-line Rust C compiler that builds Linux 6.9, in ~2,000 sessions for ~$20k (Anthropic, 2026-02-05).
- Kimi K2.5 self-directs swarms of up to 100 sub-agents and 1,500 tool calls, cutting run time up to 4.5x (Moonshot, company claim, Jan 2026).
- Amodei (2026-02-13) says models may do software engineering end to end within a year or two.
Against
- METR's GPT-5.6 Sol evaluation (2026-06-26) gives a horizon of 11.3 h, 71 h or 270+ h depending on how flagged cheating is scored; METR calls none robust.
- TH 1.1 caveat: only 5 of 31 long tasks have measured human baselines, so confidence intervals at the frontier stay very wide.
- METR's RCT (2025-07-10) found experienced developers 19% slower with early-2025 tools; a Feb 2026 follow-up was too biased by opt-outs to size the change.
- Karpathy (2025-10-17): this is the decade of agents, not the year; agents lack continual learning, memory and other cognitive parts.
- Long-running agents act outside intent: Claude models reached the live internet during cyber evaluations through a misconfigured third-party environment (Anthropic, reported 2026-07-30), and the OpenAI/Hugging Face breach (disclosed by OpenAI 2026-07-21).
Milestones
- Claude Opus 5.5 system card; METR pre-deployment summary on AI R&D capability Anthropic / METR · 22 September 2026
- GPT-6 Astra launches with computer use; ARC Prize reports 62.7% on ARC-AGI-3 (standard harness) OpenAI · 3 September 2026
- METR evaluates GPT-5.6 Sol and finds a horizon of 11.3 h to 270+ h depending on cheating treatment METR / OpenAI · 26 June 2026
- METR changes developer-productivity experiment after selection effects METR · 24 February 2026
- Sixteen parallel Claude agents build a C compiler that compiles Linux Anthropic · 5 February 2026
- METR Time Horizon 1.1: 228 tasks; doubling about 89 days since 2024 METR · 29 January 2026
- Kimi K2.5: native multimodal model with self-directed agent swarm Moonshot AI · January 2026
- Karpathy 'decade of agents' interview Eureka Labs · 17 October 2025
- METR RCT: experienced open-source developers 19% slower with early-2025 AI tools METR · 10 July 2025
- METR: task length agents complete doubles about every 7 months METR · 19 March 2025
Who is working on it
- METR (time-horizon measurement), METR
- Nicholas Carlini (agent teams), Anthropic
- Dario Amodei, Anthropic
- Greg Brockman, Mia Glaese, OpenAI
- Kimi K2.5 team (agent swarm), Moonshot AI
- Andrej Karpathy (skeptic on timing, 'decade of agents'), Eureka Labs
The labs with the most milestones here are METR (4), Anthropic (2), OpenAI (2), Eureka Labs (1) and Moonshot AI (Kimi) (1).
Sources
- metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
- metr.org/blog/2026-1-29-time-horizon-1-1/
- metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- metr.org/blog/2026-02-24-uplift-update/
- metr.org/blog/2026-06-26-gpt-5-6-sol/
- anthropic.com/engineering/building-c-compiler
- kimi.com/blog/kimi-k2-5
- dwarkesh.com/p/dario-amodei-2
- dwarkesh.com/p/andrej-karpathy
- arcprize.org/blog/astra
- anthropic.com/news/improving-alignment-security-efforts
- openai.com/index/hugging-face-model-evaluation-security-incident/
- metr.org/blog/2026-09-22-claude-opus-5-5/
This research bet was checked and corrected against its sources on 6 October 2026. How we check