Program synthesis, neuro-symbolic systems and ARC-style generalization

Combine neural intuition with discrete program search or refinement loops to learn new tasks from few examples; ARC Prize tracks progress.

Pattern interpolation generalizes poorly to novel tasks. Chollet's Ndea bets program synthesis, guided by deep learning, is the missing paradigm; ARC Prize 2025 named iterative 'refinement loops' the year's theme. ARC-AGI-3 now tests interactive exploration, modeling, goal-setting and planning.

Where it stands. ARC-AGI-3 went from frontier 0.51% (March) to 62.7% standard and 99.9% adapter-harness for GPT-6 Astra (September), at $19-26k a run. ARC Prize 2026 ($2M+) is open.

Evidence

For

  • ARC Prize 2025 (2025-12-05): 1,455 teams; top Kaggle score 24% on ARC-AGI-2 from NVARC, combining test-time-trained models with TRM components.
  • Poetiq's Gemini 3 Pro refinement harness scored 54% on ARC-AGI-2 versus 37.6% for Claude Opus 4.5 Thinking (ARC Prize, Dec 2025).
  • CompressARC, 76K parameters trained only at test time with no pretraining, won third paper prize in 2025; TRM won first.
  • GPT-6 Astra (2026-09-03): 62.7% on ARC-AGI-3 semi-private ($26k, standard harness) and 99.9% ($19k, provider harness); fewer actions than the median human on 96% of levels.
  • Astra converted unseen game mechanics into compact symbolic notes it invented, a program-like world model built in context (ARC Prize).

Against

  • ARC Prize says saturating ARC-AGI-3 would not prove AGI: the games are deterministic, closed-ended and tightly bounded.
  • Inference: the jump from 0.43% (GPT-5.5, 2026-05-01) to 62.7% (Astra) came from a general frontier LLM with large pretraining and RL, not from program-synthesis systems.
  • Results depend on harness and cost: 62.7% vs 99.9% across harnesses, at $19-26k per full run, while ARC pays humans ~$12.78 per attempted game.
  • Sources differ at launch: Fortune cites 66% on the standard harness, while ARC Prize reports 62.7% on its semi-private set (2026-09-03); OpenAI's own launch page gives only the 99.9% figure.
  • HRM's ARC result came mostly from an outer refinement loop and task-specific training, not its hierarchical design (ARC Prize, 2025-08-15).
  • Not found: a frontier lab shipping a program-synthesis system as its main product; Ndea has published no frontier results.

Milestones

Who is working on it

  • Francois Chollet, Ndea / ARC Prize Foundation
  • Greg Kamradt, ARC Prize Foundation
  • Poetiq team (refinement harnesses), Poetiq
  • Gary Marcus (neurosymbolic advocate), Marcus on AI
  • Alexia Jolicoeur-Martineau (TRM), Samsung SAIL Montreal

The labs with the most milestones here are ARC Prize Foundation (6), Samsung SAIL Montreal (1), Sapient Intelligence (1) and OpenAI (1).

Sources

  1. ndea.com/
  2. arcprize.org/blog/arc-prize-2025-results-analysis
  3. arcprize.org/blog/arc-agi-3-launch
  4. arcprize.org/blog/arc-agi-3-gpt-5-5-opus-4-7-analysis
  5. arcprize.org/blog/astra
  6. fortune.com/2026/09/03/openai-debuts-gpt-6-astra-computer-use-greg-brockman-says-start-of-
  7. arcprize.org/blog/hrm-analysis
  8. arxiv.org/abs/2510.04871
  9. arxiv.org/abs/2506.21734
  10. garymarcus.substack.com/p/a-trillion-dollars-is-a-terrible
  11. openai.com/index/gpt-6-astra/

This research bet was checked and corrected against its sources on 6 October 2026. How we check