Geiping et al. propose a language model that reasons by looping a recurrent block

"Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach" (arXiv v1 2025-02-07, nine authors) proposed a language model that iterates a recurrent block, unrolled to arbitrary…

Date
7 February 2025
Who
academic collaboration led by Jonas Geiping
Confidence
High
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): academic collaboration led by Jonas Geiping · Confidence: High "Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach" (arXiv v1 2025-02-07, nine authors) proposed a language model that iterates a recurrent block, unrolled to arbitrary depth at test time, so that it scales test-time compute by reasoning in latent space instead of emitting more tokens. It needs no specialised chain-of-thought data, and a proof-of-concept 3.5B-parameter model trained on 800B tokens improved on reasoning benchmarks, sometimes dramatically, up to a computation load equivalent to 50B parameters (arXiv 2502.05171; the weights are public as Huginn-0125). It is the reference point for the "recurrent depth" label later attached to GPT-6 Astra and shows the idea is not new to OpenAI. Related work includes Universal Transformers (2018), Mixture-of-Recursions (arXiv 2507.10524, 2025-07-14), ByteDance's Ouro looped language models (arXiv 2510.25741, 2025-10-29) and the small recursive reasoners of B08-31a (Raschka's survey lists these and more). Safety researchers care because a model that does more computation per emitted token gives a monitor less text to read (B08-33). Sources: arXiv 2502.05171 · arXiv 2507.10524 · arXiv 2510.25741

Read it in the deep dive