DeepSeek's $5.576M for V3 and $294K for R1 leave out research, failed runs and the cluster

DeepSeek disclosed $5.576M of GPU rental for V3's final run and a further $294K for R1's reasoning RL, and neither figure includes the research, the failed runs or the cluster.

Date
31 January 2025
Who
DeepSeek, SemiAnalysis, Epoch AI, Nature
People
Nathan Lambert, SemiAnalysis, Dario Amodei
Confidence
High on the disclosed figures; Medium on third-party fleet estimates
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Landmark · Significance: 4/5 · Org(s): DeepSeek, SemiAnalysis, Epoch AI, Nature · People: Nathan Lambert, SemiAnalysis, Dario Amodei · Confidence: High on the disclosed figures; Medium on third-party fleet estimates Primary sources: V3 report · R1 paper v2, Table 7 · SemiAnalysis, "DeepSeek Debates" (2025-01-31) · Nathan Lambert, Interconnects

One-liner. DeepSeek disclosed $5.576M of GPU rental for V3's final run and a further $294K for R1's reasoning RL, and neither figure includes the research, the failed runs or the cluster.

What DeepSeek disclosed.

FigureWhat it coversWhat it excludesSource
$5.576M (2.788M H800 GPU-hours at $2/hour)V3 final pre-training, context extension and post-training, on a 2,048-H800 clusterPrior research and ablations (stated); data, staff, hardware purchaseV3 paper
$294K (147K H800 GPU-hours, split as R1-Zero 101K, SFT-data creation 5K, R1 41K)Reasoning RL on 512 H800s, running about 198 hours for R1-Zero and about 80 hours for R1The V3-Base pre-train; experiments on A100s with a 30B model; data and staffR1 v2 paper, App. B.4.4 and Table 7 (published in Nature, 2025-09-17)
Combined $5.87MV3 final run plus R1 RLSame exclusionsSum of the two (my inference); The Register's version is "closer to $5.87 million" (The Register)

What others estimated. SemiAnalysis (2025-01-31) put DeepSeek's fleet at about 50,000 Hopper-generation GPUs (roughly 10,000 H800s and 10,000 H100s plus H20 orders, shared with parent High-Flyer), with total server capex near $1.6B and about $944M of operating cost, and noted the $6M excludes R&D, total cost of ownership and months of architecture work such as MLA (SemiAnalysis). Dario Amodei cited about 50,000 chips worth roughly $1B (essay). Nathan Lambert estimated the experiments behind a final run at 2-4x the reported number and the full organisation at about $500M a year (or $1B+ if operating in the US) (Interconnects). Lambert also stressed that the efficiency is real on its own terms. He compares V3's pre-training figure (2.664M GPU-hours, which he rounds to 2.6M) with 30.8M for Llama 3.1 405B, roughly one-tenth; DeepSeek's own all-in 2.788M gives 11.0x and 2.664M gives 11.6x. Epoch AI, before the Nature disclosure, estimated R1's RL stage at roughly 20% of V3's pre-training cost (about $1M against about $5M), well above the $294K DeepSeek later reported (Epoch, 2025-05).

Inference economics. On 2025-03-01 DeepSeek published a one-day snapshot of its serving system. At $2 per H800-hour the cost was $87,072 a day against $562,027 of "theoretical" revenue at list prices, a 545% theoretical margin, while saying actual revenue was far lower because of free web/app use, off-peak discounts and cheaper V3 traffic (reported by ARY; primary DeepSeek post not retrieved).

Why it matters. Three claims need to be kept apart. (1) Cheap relative to peers per capability: supported, driven by sparse MoE, FP8 and communication-aware engineering (B02). (2) Cheap in total: not supported; fleet and R&D are in the hundreds of millions to billions. (3) Reasoning was cheap to add to a good base in early 2025. I infer that the $294K covers incremental RL on an existing base and excludes the base, ablations and pilots, so it shows the recipe was short. It does not show that every lab could replicate it in months for that sum, and the lags in the Diffusion map point to RL infrastructure and engineering capacity as the pacing item (B08-10). The cheapness did not persist at the frontier (V3.2 put RL above 10% of pre-training cost). The v2 paper also says DeepSeek used A100s for small-model pilots, and the dollar figures use rented-GPU pricing, which is a different basis from owned-hardware accounting.

Nuance and myths. "R1 cost $294K" is wrong as a total; it is incremental RL on an existing base ("closer to $5.87M" with the base's final run, still excluding everything else). OpenAI-style frontier runs are priced by amortised cluster cost, not rental, so the numbers are not like-for-like. DeepSeek's RL spend also grew after R1, and its later report says the RL stage in V3.2 exceeded 10% of pre-training cost (B08-43).

Interview kit.

  • 30-second version: DeepSeek disclosed $5.6M for V3's last training run and $294K for R1's RL on top; both exclude R&D, failed runs and the roughly 50,000-GPU fleet that analysts estimate cost $1B-1.6B.
  • Likely follow-ups: So was it 100x cheaper? → Per final run, V3 used about 11 to 12 times fewer GPU-hours than Llama 3.1 405B (11.0x with DeepSeek's 2.788M, 11.6x with Lambert's 2.664M pre-training figure); counting everything, no. Why was RL so cheap? → 147K GPU-hours of RL on a ready base. Did the efficiency matter? → Yes, but it raised demand for compute.
  • Common mistake: Presenting $294K as "the cost of R1".
  • Connect it to: B02, B17, B18.

Sources. All opened except the Nature page (login redirect), Bloomberg (403) and DeepSeek's 2025-03-01 post (secondary only).

Read it in the deep dive