Small and on-device models
Capable 1-30B-parameter models that run locally or cheaply, via distillation, MoE and quantization, argued to be the right workhorse for most agent calls.
Most agent calls are narrow and repetitive, so specialised small models are cheaper and faster than frontier LLMs (NVIDIA position paper). Karpathy envisions a ~1B 'cognitive core' with reasoning but little memorised trivia; Google, Apple and Liquid ship on-device models.
Where it stands. Small models are mainstream for on-device and sub-agent use but do not lead frontier benchmarks; the 'cognitive core' idea is unproven. Gemma 4 E2B/E4B (Apr 2026) are Google's current on-device open models.
Evidence
For
- NVIDIA researchers argue small language models are powerful enough, more suitable and cheaper for most agentic subtasks (2025-06-02; v3 revised 2026-09-22).
- Gemma 4 (2026-04-02): E2B and E4B on-device sizes plus 26B MoE and 31B dense, Apache 2.0; the 31B ranks #3 open model on Arena AI's text leaderboard (Google).
- Meta's Muse Glimmer (2026-08-10): a 30B-parameter Apache 2.0 open model for always-on local agents, sized for a single consumer GPU (company-described).
- Apple's 2025 report describes a ~3B on-device model using KV-cache sharing and 2-bit quantization-aware training.
- Liquid AI's LFM2 (2025-07-10): hybrid architecture with 2x faster decode and prefill than Qwen3 on CPU (company claim).
- TRM: a 7M-parameter model beats most LLMs on ARC-AGI-1 puzzles (2025-10-06).
Against
- Not found: a small model leading frontier agent or long-horizon benchmarks; those are led by the largest models (GPT-6 Astra, Gemini 4 Argon, Claude Opus 5.x).
- Karpathy's 'cognitive core' is a proposal; no ~1B model that reasons well without stored knowledge has been shown (not found).
- Google positions Gemma as a complement to its Gemini models, not a replacement (2026-04-02).
- NVIDIA's argument is a position paper, not a measured benchmark of small versus large agents.
Milestones
- Muse Glimmer is an open 30B agentic model that runs on a single consumer GPU Meta Superintelligence Labs · 10 August 2026
- Gemma 4: E2B, E4B, 26B MoE and 31B dense under Apache 2.0 Google DeepMind · 2 April 2026
- Karpathy proposes a ~1B-parameter 'cognitive core' Eureka Labs · 17 October 2025
- Tiny Recursive Model has 7M parameters and scores 45% on ARC-AGI-1 Samsung SAIL Montreal · 6 October 2025
- LFM2: fastest on-device foundation models (company claim) Liquid AI · 10 July 2025
- Apple Intelligence Foundation Language Models Tech Report 2025: ~3B on-device model Apple · July 2025
- Gemma 3n is released in full as mobile-first multimodal open models Google DeepMind · 26 June 2025
- 'Small Language Models are the Future of Agentic AI' (arXiv v1) NVIDIA Research · 2 June 2025
Who is working on it
- Peter Belcak, Pavlo Molchanov, NVIDIA Research
- Andrej Karpathy, Eureka Labs
- Gemma team, Google DeepMind
- Foundation Models team, Apple
- LFM team, Liquid AI
The labs with the most milestones here are Google DeepMind (2), Apple (1), Liquid AI (1), NVIDIA Research (1), Samsung SAIL Montreal (1) and Eureka Labs (1).
Sources
- arxiv.org/abs/2506.02153
- blog.google/innovation-and-ai/technology/developers-tools/gemma-4/
- developers.googleblog.com/en/introducing-gemma-3n-developer-guide/
- machinelearning.apple.com/research/apple-foundation-models-tech-report-2025
- liquid.ai/blog/liquid-foundation-models-v2-our-second-series-of-generative-ai-models
- dwarkesh.com/p/andrej-karpathy
- research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model
This research bet was checked against its sources on 6 October 2026. How we check