Small and on-device models

Capable 1-30B-parameter models that run locally or cheaply, via distillation, MoE and quantization, argued to be the right workhorse for most agent calls.

Most agent calls are narrow and repetitive, so specialised small models are cheaper and faster than frontier LLMs (NVIDIA position paper). Karpathy envisions a ~1B 'cognitive core' with reasoning but little memorised trivia; Google, Apple and Liquid ship on-device models.

Where it stands. Small models are mainstream for on-device and sub-agent use but do not lead frontier benchmarks; the 'cognitive core' idea is unproven. Gemma 4 E2B/E4B (Apr 2026) are Google's current on-device open models.

Evidence

For

  • NVIDIA researchers argue small language models are powerful enough, more suitable and cheaper for most agentic subtasks (2025-06-02; v3 revised 2026-09-22).
  • Gemma 4 (2026-04-02): E2B and E4B on-device sizes plus 26B MoE and 31B dense, Apache 2.0; the 31B ranks #3 open model on Arena AI's text leaderboard (Google).
  • Meta's Muse Glimmer (2026-08-10): a 30B-parameter Apache 2.0 open model for always-on local agents, sized for a single consumer GPU (company-described).
  • Apple's 2025 report describes a ~3B on-device model using KV-cache sharing and 2-bit quantization-aware training.
  • Liquid AI's LFM2 (2025-07-10): hybrid architecture with 2x faster decode and prefill than Qwen3 on CPU (company claim).
  • TRM: a 7M-parameter model beats most LLMs on ARC-AGI-1 puzzles (2025-10-06).

Against

  • Not found: a small model leading frontier agent or long-horizon benchmarks; those are led by the largest models (GPT-6 Astra, Gemini 4 Argon, Claude Opus 5.x).
  • Karpathy's 'cognitive core' is a proposal; no ~1B model that reasons well without stored knowledge has been shown (not found).
  • Google positions Gemma as a complement to its Gemini models, not a replacement (2026-04-02).
  • NVIDIA's argument is a position paper, not a measured benchmark of small versus large agents.

Milestones

Who is working on it

  • Peter Belcak, Pavlo Molchanov, NVIDIA Research
  • Andrej Karpathy, Eureka Labs
  • Gemma team, Google DeepMind
  • Foundation Models team, Apple
  • LFM team, Liquid AI

The labs with the most milestones here are Google DeepMind (2), Apple (1), Liquid AI (1), NVIDIA Research (1), Samsung SAIL Montreal (1) and Eureka Labs (1).

Sources

  1. arxiv.org/abs/2506.02153
  2. blog.google/innovation-and-ai/technology/developers-tools/gemma-4/
  3. developers.googleblog.com/en/introducing-gemma-3n-developer-guide/
  4. machinelearning.apple.com/research/apple-foundation-models-tech-report-2025
  5. liquid.ai/blog/liquid-foundation-models-v2-our-second-series-of-generative-ai-models
  6. dwarkesh.com/p/andrej-karpathy
  7. research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model

This research bet was checked against its sources on 6 October 2026. How we check