DeepSeek released V3, the 671B open base model that R1 was trained on
A 671B-parameter open mixture-of-experts model matched the best closed non-reasoning models of its day for a reported 2.788M H800 GPU-hours, and became the base on which R1 was trained.
- Date
- 26 December 2024
- Who
- DeepSeek
- People
- Liang Wenfeng (founder), the DeepSeek-AI team
- Confidence
- High (paper figures); the "$5.6M" interpretation is contested, see B08-15
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Landmark · Significance: 4/5 · Org(s): DeepSeek · People: Liang Wenfeng (founder), the DeepSeek-AI team · Confidence: High (paper figures); the "$5.6M" interpretation is contested, see B08-15 Primary sources: DeepSeek-V3 Technical Report, arXiv 2412.19437 (v1 2024-12-27, v2 2025-02-18) · DeepSeek API news, 2024-12-26
One-liner. A 671B-parameter open mixture-of-experts model matched the best closed non-reasoning models of its day for a reported 2.788M H800 GPU-hours, and became the base on which R1 was trained.
Why it happened. DeepSeek grew out of the AI research of the Chinese quant fund High-Flyer and was launched in 2023 under its founder Liang Wenfeng; High-Flyer had begun buying large GPU fleets in 2021 (Fortune profile; lab history in B20). US export controls capped what chips it could buy (the cluster used export-compliant H800s). I infer that this pushed the lab to co-design model, training software and cluster around scarce interconnect bandwidth. The report is explicit that cross-node InfiniBand is about 50 GB/s versus about 160 GB/s of intra-node NVLink, and that the routing algorithm was built to fit that. V3 extended DeepSeek-V2's design (May 2024) with Multi-head Latent Attention (MLA) to shrink the key-value (KV) cache of stored attention state and DeepSeekMoE for sparse activation (B02).
The idea. Make each token cheap, then train on a lot of them. How it works: 671B total parameters with 37B active per token; 256 routed experts, 8 chosen per token; trained on 14.8T tokens at a 4K sequence length, then extended to 32K and 128K. New pieces were an auxiliary-loss-free load-balancing method (no extra loss term to fight expert collapse), a multi-token-prediction training objective, FP8 mixed-precision training (8-bit floating-point arithmetic) that the authors say is the first validated at this scale, and "DualPipe" pipeline parallelism that hides communication behind computation; only 20 streaming multiprocessors per GPU were needed for cross-node communication. The report says pretraining had no irrecoverable loss spikes and no rollbacks. In post-training it distilled reasoning behaviour from an internal R1-series long-CoT model into the standard chat model (see R1-Lite).
Results. MMLU 88.5, MMLU-Pro 75.9, GPQA Diamond 59.1; on maths it beat o1-preview on MATH-500 and was the top non-long-CoT model on several maths and code benchmarks; the authors place it at GPT-4o / Claude 3.5 Sonnet level on general tasks. Released with weights on 2024-12-26 at API prices of $0.27 / $1.10 per million input / output tokens after a promotional period (company page).
The cost claim. Pre-training took 2,664K H800 GPU-hours, context extension 119K and post-training 5K, 2,788K in total, which at an assumed $2 per GPU-hour rental is $5.576M. The paper states this covers only the "official training run" and excludes prior research and ablation experiments on architectures, algorithms and data (report, §1). The whole debate about "DeepSeek cost $6M" is a debate about that sentence.
How it spread. The efficiency ideas travelled by paper and open weights. DeepSeek's own later models extend these ideas (sparse attention in V3.2, compressed attention in V4); whether other labs adopted MLA, FP8 training or multi-token prediction is tracked in B02 and B19 and was not verified here. More important for this chapter, V3-Base was the fertile pre-trained model that made R1-Zero a roughly $200K RL run rather than a pretraining-scale one.
Why it mattered. Anthropic's Dario Amodei, writing after R1, argued that V3 was the real engineering achievement and R1 the lesser one, and that V3 was still an expected point on an ongoing cost-reduction curve. By his account DeepSeek was roughly 7-10 months behind US models, with algorithmic efficiency improving about 4x a year, and held on the order of 50,000 Hopper-generation chips worth about $1B (essay, paraphrased; compare B18).
Nuance, controversy and myths. The R1 paper's background section notes that the web-crawled pre-training data of V3-Base may contain OpenAI-model outputs, which DeepSeek says it did not add intentionally (R1 paper, App. A.1). "Trained for $6M" is wrong as a statement of total cost, as the cost entry explains.
Interview kit.
- 30-second version: V3 is a very sparse 671B MoE with FP8 training and attention compression that cost about $5.6M in GPU rental for its final run; it is the base model under R1.
- Likely follow-ups: Why cheap? → Only 37B active parameters per token, low-precision training, communication hidden behind compute. What is excluded? → R&D, ablations, data, salaries, the cluster itself.
- Common mistake: Saying R1 cost $6M; the $5.6M is V3's final pre-train, and R1's RL cost is a separate $294K figure.
- Connect it to: B02, B03, B20.
Sources. Paper and API page above (opened; paper read locally); Amodei essay (opened).