Meta releases MobileLLM-R1, its first released reasoning models
Meta released MobileLLM-R1 on 2025-09-12 (365 days after o1-preview, 235 after R1).
- Date
- 12 September 2025
- Who
- Meta (AI at Meta)
- Confidence
- High on release facts (model card and arXiv abstract opened)
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 2/5 · Org(s): Meta (AI at Meta) · Confidence: High on release facts (model card and arXiv abstract opened) Meta released MobileLLM-R1 on 2025-09-12 (365 days after o1-preview, 235 after R1). The release is three sub-billion models (140M, 360M and 950M parameters) specialised for maths, Python and C++ coding and scientific problems and not aimed at general chat. Per the model card, the 950M model was pre-trained on about 2T curated tokens (4.2T resampled tokens in the paper's accounting), mid-trained with knowledge distillation from Llama-3.1-8B, and post-trained with supervised fine-tuning on reasoning data; the card mentions no RL. The paper (arXiv 2509.24945, submitted 2025-09-29, ICLR 2026) reports 15.5 on AIME for the 950M model and says it matches or beats Qwen3-0.6B using 11.7% of Qwen3's pre-training tokens; the licence is FAIR Noncommercial Research (model card). These are Meta's first released reasoning models. That is why Muse Spark is "Meta's first flagship reasoning model," not its first reasoning model, and why Meta's lag has two values in the Diffusion map (365 days for the small SFT-only models, 573 for the flagship). Sources: model card · arXiv 2509.24945