DeepSeek
Efficiency-first open frontier models, with sparse and compressed attention for 1M-token context at a fraction of compute, MoE at 1.6T parameters, aggressive pricing and optimization for Chinese chips. A small, low-hierarchy research team with no KPIs, per founder.
- Founded
- July 2023
- Headquarters
- Hangzhou, China
- Founders
- Liang Wenfeng
- Pushes against
- Spending-driven frontier scaling; closed weights.
- Now
- V4 matches mid-2026 frontier-minus-one models on company-reported benchmarks. Closed a first ~$7B round (2026-06), then paused a second after leaked founder remarks that China still lags and depends on Nvidia (Bloomberg, 2026-07). IPO prep reported.
Key people
- Liang Wenfeng (founder, CEO)
Funding
- 6 October 2026 reported
- Amount
- Reported at least 80 billion yuan, about $12 billion, possibly up to 100 billion yuan, about $15 billion
- Valuation
- About 500 billion yuan, about $74 to 75 billion, possibly higher
- Lead
- CATL and Tencent with the largest commitments, plus Geely Auto, Monolith Management and Loyal Valley Capital
- July 2026 reported
- Amount
- Second round, paused
- Valuation
- ~$71B pre-money (reported)
- June 2026 reported
- Amount
- ~50B yuan (~$7.4B), first outside round
- Valuation
- $52B-$59B post-money (reported)
- Lead
- Founder Liang (20B yuan), Tencent, CATL; national AI fund and others (reported)
Every funding round in the atlas
Latest launches
- DeepEP-Ascend and DeepGEMM-Ascend 29 September 2026
DeepSeek ports its MoE communication and FP8 GEMM libraries to Huawei Ascend NPUs, targeting Ascend 950 with CANN 9.2.0. - DeepSeek Elastic Compute (DSec) sandbox platform 19 September 2026
DeepSeek describes the sandbox platform behind its agentic RL: about 3 million sandboxes a day, 380,000+ concurrent, 5,000+ created per second. - DeepSelect TopK kernels and DeepJIT runtime 10 September 2026
DeepSeek opens its sparse-attention TopK kernel (2-20x faster than torch.topk) and a lightweight CUDA and Ascend JIT runtime. - DeepSeek-V4.1-Flash 10 September 2026
Native-multimodal 552B-backbone model with a causal encoder-decoder (8B active in prefill, 16B in decode) and a KV cache of 890 bytes per token. - DeepSeek-V4-Flash-Vision-Exp 21 August 2026
First multimodal V4 model, an experimental vision add-on to V4-Flash that matches its text ability; ships with a new Files API. - DeepSeek Harness (dsh) 13 August 2026
DeepSeek's own open-source agent harness ('Everything is a Plugin') with web, desktop and SSH UI, plugins, subagents and MCP. It is a developer preview under MIT. - DeepSeek-V4-Pro-0813 (general availability) 13 August 2026
V4-Pro goes GA with large agentic gains (Terminal Bench 2.1: 87.9, DeepSWE 62.7), low/high/max effort, Responses API and 50% off-peak pricing. - DeepSeek-V4-Flash-0731 (official) 31 July 2026
Official V4-Flash, with agent benchmarks up sharply versus the preview (DeepSWE 12.8 to 54.4, Terminal Bench 2.1 72.1 to 82.7) and a native Responses API for Codex.
Launches by year
67 launches and papers since 1 November 2023, oldest first within each year.
- 2026 16
- DeepSeek-R1 paper v2 (expanded on arXiv), Engram: Conditional Memory via Scalable Lookup, Engram: conditional memory via scalable lookup, DeepSeek-OCR 2 (Visual Causal Flow), TileKernels (TileLang kernel library), DeepSeek-V4 (CSA/HCA attention, mHC, Muon), DeepSeek-V4-Pro and V4-Flash (preview), DSpark speculative decoding and DeepSpec, DeepSeek-V4-Flash-0731 (official), DeepSeek-V4-Pro-0813 (general availability), DeepSeek Harness (dsh), DeepSeek-V4-Flash-Vision-Exp, DeepSeek-V4.1-Flash, DeepSelect TopK kernels and DeepJIT runtime, DeepSeek Elastic Compute (DSec) sandbox platform, DeepEP-Ascend and DeepGEMM-Ascend
- 2025 28
- DeepSeek-R1 and R1-Zero, DeepSeek-R1: Incentivizing Reasoning via RL (arXiv), Janus-Pro (1B, 7B), DeepSeek chat app tops US App Store, Native Sparse Attention (NSA), Native Sparse Attention (NSA), Open Source Week (FlashMLA, DeepEP, DeepGEMM, DualPipe, 3FS), DeepSeek-V3/R1 inference system overview ('one more thing'), DeepSeek-V3-0324, DeepSeek-GRM / SPCT (inference-time scaling for reward models), DeepSeek-Prover-V2 (7B, 671B), Insights into DeepSeek-V3: hardware co-design paper, DeepSeek-R1-0528, DeepSeek-R1-0528-Qwen3-8B, DeepSeek-V3.1, DeepSeek-R1 in Nature (peer-reviewed), DeepSeek-R1 published in Nature, DeepSeek-V3.1-Terminus, DeepSeek Sparse Attention (V3.2-Exp), DeepSeek-V3.2-Exp (DeepSeek Sparse Attention), DeepSeek-OCR (Contexts Optical Compression), DeepSeek-OCR (contexts optical compression), LPLB: linear-programming expert-parallel load balancer, DeepSeekMath-V2, DeepSeek-V3.2 and V3.2-Speciale, DeepSeek-V3.2 (DSA, scaled RL, agentic synthesis), mHC: Manifold-Constrained Hyper-Connections, mHC: Manifold-Constrained Hyper-Connections
- 2024 21
- DeepSeekMoE, DeepSeekMoE, DeepSeekMath (introduces GRPO), DeepSeekMath 7B (introduces GRPO), DeepSeek-VL (1.3B, 7B), DeepSeek-V2, DeepSeek-V2 (Multi-head Latent Attention), DeepSeek-Prover (V1), DeepSeek-Coder-V2, ESFT: Expert-Specialized Fine-Tuning for MoE, API context caching on disk, DeepSeek-Prover-V1.5, Fire-Flyer AI-HPC: cost-effective cluster co-design, Auxiliary-loss-free load balancing for MoE, DeepSeek-V2.5, Janus (and JanusFlow), DeepSeek-R1-Lite-Preview, DeepSeek-V2.5-1210, DeepSeek-VL2, DeepSeek-V3, DeepSeek-V3 Technical Report
- 2023 2
- DeepSeek Coder (1.3B-33B), DeepSeek LLM 7B / 67B
People moves
Of the moves the atlas tracks, 0 joined and 1 left.
- Luo Fuli 12 November 2025
DeepSeek to Xiaomi
Core developer, DeepSeek
Sources
- api-docs.deepseek.com/news/news260424
- api-docs.deepseek.com/news/news260910
- cnbc.com/2026/06/03/deepseek-slated-to-draw-7-billion-in-maiden-fundraising-sources-say.ht
- finance.yahoo.com/technology/ai/articles/deepseek-hits-brakes-mega-funding-131502562.html
- technologyreview.com/2026/04/24/1136422/why-deepseeks-v4-matters/
- fortune.com/2026/08/01/deepseek-founder-liang-wenfeng-workers-dont-do-overtime-or-kpis-chi
- en.wikipedia.org/wiki/DeepSeek
This profile was checked against its sources on 6 October 2026. How we check