Windows local open-weight models and llama.cpp in Windows ML
Microsoft announced at its Windows keynote that DeepSeek V4 Flash and a new Nvidia Nemotron will run locally on RTX Spark PCs, and that llama.cpp is coming to Windows ML.
Microsoft is positioning Windows as a platform for hybrid intelligence, where work runs on the PC or in the cloud depending on the task. Heavily compressed large models are brought to the PC, including DeepSeek V4 Flash (284 billion parameters, quantized to 1.6 bits per the keynote slide) and an upcoming Nemotron of more than 70 billion parameters. Developers can use open-source models through llama.cpp in Windows ML, and GitHub HydraFusion will route tasks to local models.
- Date
- Wednesday 7 October 2026
- Lab
- Microsoft
- Kind
- feature
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| DeepSeek V4 Flash parameters | 284 billion quantized to 1.6 bits, runs in 60 GB of memory, read off keynote slides | company |
| Nemotron local model parameters | more than 70 billion about 20 GB of memory at 2-bit precision, expected October 15 | company |
| MAI Code 1.1 Flash parameters | 137 billion (6.8 billion active) 3-bit precision, 256,000 token context | company |
Figures come from keynote slides as reported by coverage, and the official Windows Blog and Nvidia posts were not yet online.
Sources
- pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-m
- trendingtopics.eu/windows-hybrid-intelligence-deepseek-nemotron/
This record was checked and corrected against its sources on 8 October 2026. How we check