Meta releases proprietary Muse Spark, its first flagship reasoning model, 573 days after o1-preview
Meta's first flagship reasoning model came 573 days after o1-preview and 443 after R1. Counting Meta's earlier small MobileLLM-R1 models (B08-37a), the lags are 365 and 235 days.
- Date
- 8 April 2026
- Who
- Meta Superintelligence Labs
- People
- Alexandr Wang
- Confidence
- Medium (company-claimed benchmarks; limited independent data)
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): Meta Superintelligence Labs · People: Alexandr Wang · Confidence: Medium (company-claimed benchmarks; limited independent data) Meta's first flagship reasoning model came 573 days after o1-preview and 443 after R1. Counting Meta's earlier small MobileLLM-R1 models (B08-37a), the lags are 365 and 235 days. Llama 4 (2025-04-05) shipped with no reasoning model (per TechCrunch, none of the Llama 4 models is a proper reasoning model; some coverage reports Meta teasing a "Llama 4 Reasoning," which I did not confirm or find released). Meta then reorganised under Superintelligence Labs and hired o1-era researchers from OpenAI including Shengjia Zhao (named chief scientist, 2025-07-25, per TechCrunch), Jason Wei and Hyung Won Chung (The Decoder, 2025-07-16), plus Hongyu Ren and Trapit Bansal per TechCrunch; all five appear in the o1 foundational-contributor list.
Muse Spark is proprietary (a break with the open-weights Llama strategy, though Meta also released the open-weight 30B Muse Glimmer under Apache 2.0 on 2026-08-10) and natively multimodal, with Instant, Thinking and "Contemplating" (parallel sub-agent) modes. Meta says a "thought compression" RL penalty on thinking time let it match peers while using far fewer reasoning tokens (58 million output tokens against 120 to 157 million for competitors on one evaluation). Humanity's Last Exam scores for Muse Spark differ by outlet and condition (VentureBeat gives 42.8 without tools and 50.4 with tools, and Decrypt gives 58% in Contemplating mode; I could not reconcile them, so quote only from Meta's own table). It also scored a weak 42.5 on ARC-AGI-2 versus mid-70s for rivals (VentureBeat, secondary benchmark roundup).
My inference is that it is the clearest case in this chapter of reasoning know-how moving by talent hire. The hires are sourced, but no source links them to Muse Spark's reasoning pipeline.
Meta has since shipped Muse Spark 1.1 (2026-07-09, Meta AI blog), 1.2 (2026-08-05, with the Muse Code agent) and 1.3 (2026-09-02), the last two via Wikipedia, a pointer. In Das's IMO harness Muse Spark 1.1 scored 26/42 (B08-48). See also B19.