Subscribe

Mellum2.1

JetBrains released Mellum2.1, an open 12B mixture-of-experts coding model with 2.5B active parameters, trained mainly with reinforcement learning for agentic coding.

The architecture is the same as Mellum2, but reinforcement learning went from a short final stage to the main part of training. JetBrains ran millions of sandboxed runs across thousands of environments, and the biggest gain was in agentic coding. Multi-token prediction makes single requests about 1.6 times faster.

Date
Thursday 8 October 2026
Lab
JetBrains
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
Token throughput under heavy load vs Qwen3.5-9Balmost 2x
JetBrains says it is the fastest model in the compared group, which also included Gemma 4 E4B and Mellum2
company
Single-request speedup with multi-token predictionabout 1.6xcompany
Parameters12B total, 2.5B active
mixture of experts
company
LiveCodeBench v6 pass@182
self-reported on the Hugging Face model card, marked unverified
company

Released under Apache 2.0. GGUF builds and the MTP head for vLLM are listed as coming soon.

Sources

  1. blog.jetbrains.com/ai/2026/10/mellum2-1-gets-to-work-a-fast-open-model-for-coding-agents/
  2. huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking
  3. news.aibase.com/news/31502

This record was checked against its sources on 9 October 2026. How we check

Read the daily brief for 8 October 2026