Apple's "The Illusion of Thinking" and the rebuttals
An Apple controlled-puzzle study claimed reasoning models collapse beyond a complexity threshold, and a fast rebuttal argued much of the "collapse" came from the test design.
- Date
- 7 June 2025
- Who
- Apple Machine Learning Research
- People
- Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh, Maxwell Horton, Samy Bengio, Mehrdad Farajtabar; A. Lawsen
- Confidence
- High on what was claimed; the interpretation is contested
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Landmark · Significance: 4/5 · Org(s): Apple Machine Learning Research · People: Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh, Maxwell Horton, Samy Bengio, Mehrdad Farajtabar; A. Lawsen · Confidence: High on what was claimed; the interpretation is contested Primary sources: Apple ML Research page · arXiv 2506.06941 (v1 2025-06-07) · Lawsen's comment, arXiv 2506.09250 (2025-06-10)
One-liner. An Apple controlled-puzzle study claimed reasoning models collapse beyond a complexity threshold, and a fast rebuttal argued much of the "collapse" came from the test design.
Why it happened. Public benchmarks (AIME, GPQA) may be contaminated by training data and give little control over difficulty. Apple's group had already published "GSM-Symbolic" (arXiv 2410.05229, 2024-10) arguing that maths benchmark scores were fragile to superficial changes; the new work moved to puzzles with adjustable complexity and inspected the thinking traces as well as the answers. It appeared days before Apple's WWDC (my inference is that the timing fed the narrative that Apple, itself without a frontier reasoning model, was downplaying the technology).
The idea. Use four puzzle families (including Tower of Hanoi, checker jumping, river crossing and blocks world) where complexity is a dial (more discs, more pieces) and every move can be verified, and compare "thinking" and non-thinking pairs of models (Claude 3.7 Sonnet, DeepSeek R1/V3, o3-mini) at equal token budgets (Apple).
Results (as reported by the authors). There were three regimes. At low complexity non-thinking models matched or beat thinking ones, at medium complexity thinking helped, and at high complexity both collapsed to near-zero accuracy. Near collapse, reasoning effort (tokens spent) declined even when budget remained, and giving the algorithm in the prompt did not rescue execution ("struggle with exact computation").
The rebuttal. Lawsen's comment (v1 2025-06-10, revised 2025-06-16; the first posting reportedly listed Anthropic's Claude Opus as a co-author, while the current listing names only Lawsen) argued that (a) the Tower of Hanoi results exceed output-token limits, with models saying so in their outputs; (b) the River Crossing instances for N greater than 5 are mathematically unsolvable, yet were scored as failures; (c) the automated scorer cannot tell reasoning failure from truncation; and (d) asking for a generating function in place of every move gave high accuracy on Hanoi cases reported as failures (arXiv 2506.09250).
Why it mattered. It reframed the debate from "benchmark scores" to "what are these models doing?" and it forced better evaluation hygiene (token-limit checks, solvability checks). Its practical message, that reasoning helps in a middle band of difficulty, matches later observations about overthinking and inverse scaling (B08-06).
Nuance, controversy and myths. "Apple proved reasoning models don't reason" overstates; "the rebuttal debunked Apple" overstates too, because the rebuttal was a preliminary test on a subset of issues and was itself not peer reviewed. A defensible reading is that the models are sample-efficient on problems near their training distribution and brittle at long exact algorithmic execution, which tool use (code) fixes (B08-24).
Interview kit.
- 30-second version: Apple showed that on puzzles with dialled-up difficulty, thinking models do better than non-thinking models only in a middle band and fail entirely at high complexity; critics said some failures were token-limit and unsolvable-instance artifacts.
- Likely follow-ups: Who is right? → Both partly; the narrow claim (brittleness on long exact procedures) survives, the broad "illusion" claim does not. Why does tool use matter? → Writing code removes the need to emit every step.
- Common mistake: Attributing the rebuttal to "Claude wrote a paper"; the arXiv comment is authored by A. Lawsen.
- Connect it to: B08-25, B23, B07.
Sources. Apple page and both arXiv abstracts opened.