Mellum2.1
JetBrains released Mellum2.1, an open 12B mixture-of-experts coding model with 2.5B active parameters, trained mainly with reinforcement learning for agentic coding.
The architecture is the same as Mellum2, but reinforcement learning went from a short final stage to the main part of training. JetBrains ran millions of sandboxed runs across thousands of environments, and the biggest gain was in agentic coding. Multi-token prediction makes single requests about 1.6 times faster.
- Date
- Thursday 8 October 2026
- Lab
- JetBrains
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Token throughput under heavy load vs Qwen3.5-9B | almost 2x JetBrains says it is the fastest model in the compared group, which also included Gemma 4 E4B and Mellum2 | company |
| Single-request speedup with multi-token prediction | about 1.6x | company |
| Parameters | 12B total, 2.5B active mixture of experts | company |
| LiveCodeBench v6 pass@1 | 82 self-reported on the Hugging Face model card, marked unverified | company |
Released under Apache 2.0. GGUF builds and the MTP head for vLLM are listed as coming soon.
Sources
- blog.jetbrains.com/ai/2026/10/mellum2-1-gets-to-work-a-fast-open-model-for-coding-agents/
- huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking
- news.aibase.com/news/31502
This record was checked against its sources on 9 October 2026. How we check