Windows local open-weight models and llama.cpp in Windows ML

Microsoft announced at its Windows keynote that DeepSeek V4 Flash and a new Nvidia Nemotron will run locally on RTX Spark PCs, and that llama.cpp is coming to Windows ML.

Microsoft is positioning Windows as a platform for hybrid intelligence, where work runs on the PC or in the cloud depending on the task. Heavily compressed large models are brought to the PC, including DeepSeek V4 Flash (284 billion parameters, quantized to 1.6 bits per the keynote slide) and an upcoming Nemotron of more than 70 billion parameters. Developers can use open-source models through llama.cpp in Windows ML, and GitHub HydraFusion will route tasks to local models.

Date
Wednesday 7 October 2026
Lab
Microsoft
Kind
feature
Access
research preview

Figures

MeasureValueMeasured by
DeepSeek V4 Flash parameters284 billion
quantized to 1.6 bits, runs in 60 GB of memory, read off keynote slides
company
Nemotron local model parametersmore than 70 billion
about 20 GB of memory at 2-bit precision, expected October 15
company
MAI Code 1.1 Flash parameters137 billion (6.8 billion active)
3-bit precision, 256,000 token context
company

Figures come from keynote slides as reported by coverage, and the official Windows Blog and Nvidia posts were not yet online.

Sources

  1. pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-m
  2. trendingtopics.eu/windows-hybrid-intelligence-deepseek-nemotron/

This record was checked and corrected against its sources on 8 October 2026. How we check

Read the daily brief for 7 October 2026