Program synthesis, neuro-symbolic systems and ARC-style generalization
Combine neural intuition with discrete program search or refinement loops to learn new tasks from few examples; ARC Prize tracks progress.
Pattern interpolation generalizes poorly to novel tasks. Chollet's Ndea bets program synthesis, guided by deep learning, is the missing paradigm; ARC Prize 2025 named iterative 'refinement loops' the year's theme. ARC-AGI-3 now tests interactive exploration, modeling, goal-setting and planning.
Where it stands. ARC-AGI-3 went from frontier 0.51% (March) to 62.7% standard and 99.9% adapter-harness for GPT-6 Astra (September), at $19-26k a run. ARC Prize 2026 ($2M+) is open.
Evidence
For
- ARC Prize 2025 (2025-12-05): 1,455 teams; top Kaggle score 24% on ARC-AGI-2 from NVARC, combining test-time-trained models with TRM components.
- Poetiq's Gemini 3 Pro refinement harness scored 54% on ARC-AGI-2 versus 37.6% for Claude Opus 4.5 Thinking (ARC Prize, Dec 2025).
- CompressARC, 76K parameters trained only at test time with no pretraining, won third paper prize in 2025; TRM won first.
- GPT-6 Astra (2026-09-03): 62.7% on ARC-AGI-3 semi-private ($26k, standard harness) and 99.9% ($19k, provider harness); fewer actions than the median human on 96% of levels.
- Astra converted unseen game mechanics into compact symbolic notes it invented, a program-like world model built in context (ARC Prize).
Against
- ARC Prize says saturating ARC-AGI-3 would not prove AGI: the games are deterministic, closed-ended and tightly bounded.
- Inference: the jump from 0.43% (GPT-5.5, 2026-05-01) to 62.7% (Astra) came from a general frontier LLM with large pretraining and RL, not from program-synthesis systems.
- Results depend on harness and cost: 62.7% vs 99.9% across harnesses, at $19-26k per full run, while ARC pays humans ~$12.78 per attempted game.
- Sources differ at launch: Fortune cites 66% on the standard harness, while ARC Prize reports 62.7% on its semi-private set (2026-09-03); OpenAI's own launch page gives only the 99.9% figure.
- HRM's ARC result came mostly from an outer refinement loop and task-specific training, not its hierarchical design (ARC Prize, 2025-08-15).
- Not found: a frontier lab shipping a program-synthesis system as its main product; Ndea has published no frontier results.
Milestones
- GPT-6 Astra scores 62.7% (standard) and 99.9% (provider harness) on ARC-AGI-3 OpenAI / ARC Prize Foundation · 3 September 2026
- ARC Prize 2026 ARC-AGI-3 Milestone Prize #1 awarded ARC Prize Foundation · 6 July 2026
- ARC analysis shows GPT-5.5 scores 0.43% and Opus 4.7 0.18% on ARC-AGI-3 ARC Prize Foundation · 1 May 2026
- ARC-AGI-3 launched (interactive games); frontier AI 0.51% vs humans 100%; $2M+ prizes ARC Prize Foundation · 25 March 2026
- ARC Prize 2025 results show refinement loops as the theme, and NVARC tops Kaggle at 24% ARC Prize Foundation · 5 December 2025
- Tiny Recursive Model has 7M parameters and scores 45% on ARC-AGI-1 and 8% on ARC-AGI-2 Samsung SAIL Montreal · 6 October 2025
- ARC Prize replicates and dissects HRM: outer refinement loop drives the score ARC Prize Foundation · 15 August 2025
- Hierarchical Reasoning Model (HRM) paper, 27M parameters Sapient Intelligence · 26 June 2025
- ARC-AGI-2 and ARC Prize 2025 announced ARC Prize Foundation · 24 March 2025
Who is working on it
- Francois Chollet, Ndea / ARC Prize Foundation
- Greg Kamradt, ARC Prize Foundation
- Poetiq team (refinement harnesses), Poetiq
- Gary Marcus (neurosymbolic advocate), Marcus on AI
- Alexia Jolicoeur-Martineau (TRM), Samsung SAIL Montreal
The labs with the most milestones here are ARC Prize Foundation (6), Samsung SAIL Montreal (1), Sapient Intelligence (1) and OpenAI (1).
Sources
- ndea.com/
- arcprize.org/blog/arc-prize-2025-results-analysis
- arcprize.org/blog/arc-agi-3-launch
- arcprize.org/blog/arc-agi-3-gpt-5-5-opus-4-7-analysis
- arcprize.org/blog/astra
- fortune.com/2026/09/03/openai-debuts-gpt-6-astra-computer-use-greg-brockman-says-start-of-
- arcprize.org/blog/hrm-analysis
- arxiv.org/abs/2510.04871
- arxiv.org/abs/2506.21734
- garymarcus.substack.com/p/a-trillion-dollars-is-a-terrible
- openai.com/index/gpt-6-astra/
This research bet was checked and corrected against its sources on 6 October 2026. How we check