Insights › Market reaction

Did DeepSeek Build R1 for $6 Million? What the Papers Actually Say

Five calendar days before Nvidia's 17% Monday drop, the R1 paper appeared—but the famous $5.576 million number belonged to V3's official training compute estimate, not an all-in R1 budget.

Published in ET: Feed time in ET: Artificial Intelligence
  • DeepSeek-V3 reported 2.788 million H800 GPU-hours and valued that official training work at $5.576 million using an assumed $2 per GPU-hour.
  • The V3 paper explicitly excluded prior research and ablation experiments involving architecture, algorithms and data.
  • The R1 paper arrived on Wednesday, January 22—five calendar days before Nvidia's 17% Monday drop—and did not disclose an all-in R1 development cost.
Research chart for Did DeepSeek Build R1 for $6 Million? What the Papers Actually Say
MoveSurge retrospective research chart. Definitions, inputs and limitations appear directly below and in the dated source ledger.

Archive note: this fact check was researched and first published in August 2026. The event time marks the first R1 paper submission.

Countdown to Nvidia's 17% drop: the R1 paper was submitted on Wednesday, January 22, 2025—five calendar days before Nvidia fell 16.9% on Monday, January 27.

The sentence that dominated DeepSeek coverage was simple: a Chinese lab had built an OpenAI-class reasoning model for about $6 million. The source documents support a narrower and still important result. They do not support that all-in R1 development claim.

Where the $5.576 million figure came from

DeepSeek's V3 paper, first submitted on Friday, December 27, 2024—thirty-one calendar days before Nvidia's drop—reported 2.788 million Nvidia H800 GPU-hours across the official V3 training work. The authors applied an assumed rental price of $2 per GPU-hour and calculated $5.576 million. The table included pretraining, context extension and post-training.

The authors also added an explicit boundary: the estimate excluded costs tied to earlier research and ablation experiments on architecture, algorithms and data. It did not represent hardware purchases, staff, data acquisition, failed experiments or the full history of the High-Flyer computing operation.

What the R1 paper said

The R1 paper arrived on Wednesday, January 22, five calendar days before Nvidia's 17% Monday drop. It described reasoning training built on DeepSeek-V3-Base, including R1-Zero, cold-start data, reinforcement learning, supervised fine-tuning and distilled models. It provided benchmark and method detail. It did not state that R1's entire research and development program cost $6 million, and it did not publish a complete R1 compute ledger.

Repeated claimPrimary-source recordAssessment
DeepSeek built R1 for $6 millionThe $5.576 million estimate was published in the V3 paper; the R1 paper gave no all-in development total.False attribution.
The number covered all DeepSeek developmentThe V3 paper excluded prior research and ablation experiments.Materially misleading.
DeepSeek trained on 2,048 H800 GPUsThe V3 paper describes a 2,048-GPU cluster and reports total GPU-hours.Supported for the documented V3 training setup, not proof of the company's entire chip inventory.
DeepSeek possessed about 50,000 H100sAlexandr Wang made the claim in a Davos interview without publishing evidence.Unverified allegation.

Why the comparison moved markets

A final training-compute estimate is not comparable with a company's annual capital budget. The first measures a bounded model run under an assumed compute price. The second can include data centers, networking, power, research staff, multiple models, inference capacity and assets with multi-year lives. Weekend posts—one to three calendar days before Nvidia's Monday drop—repeatedly put the V3 estimate beside U.S. corporate spending as if both numbers described the same unit of work.

The correction does not erase DeepSeek's achievement. V3 and R1 demonstrated serious efficiency improvements, open-weight distribution and competitive performance under tighter hardware constraints than leading U.S. labs. The correction changes the market question from “Can anyone reproduce R1 from scratch for $6 million?” to “How far did DeepSeek shift the cost-performance curve after years of research and access to a substantial compute base?”

The chip claim needs the same discipline

DeepSeek documented H800 use for V3. On Thursday, January 23—four calendar days before Nvidia's drop—Wang said his understanding was that DeepSeek had about 50,000 H100 chips, and Elon Musk replied “Obviously” to a clip of that claim. Neither response supplied a verifiable inventory. High-Flyer had publicly described a 10,000-A100 cluster years earlier, which establishes meaningful prior infrastructure but does not verify the H100 allegation.

The cost and chip stories therefore contained facts, omissions and unverified claims moving in both directions. The X amplification analysis labels each high-reach post by evidence status rather than treating every viral assertion as equivalent.

How one scoped number changed meaning

The transmission chart above follows the claim rather than the share price. DeepSeek's Friday, December 27, 2024 V3 report—thirty-one calendar days before Nvidia's drop—described 2.788 million H800 GPU-hours across the official V3 training stages. At an assumed $2 per GPU-hour, the authors calculated $5.576 million. The report immediately narrowed the scope by excluding the costs of earlier research and ablation experiments on architectures, algorithms and data.

Downstream retellings progressively removed four boundaries: V3 became R1; documented training compute became total development; an assumed compute rental value became cash spent; and one model's work became the economics of the whole company. By Sunday, January 26—one calendar day before the drop—social posts were comparing the scoped number directly with annual corporate infrastructure plans. By Monday, January 27—the selloff day—some headlines described a model as having been developed in two months for under $6 million.

A claim-by-claim scope audit

Amplified formulationWhat the evidence supportsVerdict
“R1 was built for $6 million”The $5.576 million calculation belongs to the documented V3 training work. The Wednesday, January 22 R1 paper—five calendar days before the drop—publishes no all-in R1 cost.False attribution.
“DeepSeek's whole program cost $6 million”The V3 report explicitly excludes prior research and ablation experiments; it does not total hardware, staff, data, failed runs or company infrastructure.Materially misleading.
“DeepSeek was 25× cheaper”The located comparison concerns API/inference pricing and selected reasoning capability. It is not a training-cost ratio.Supported only with scope.
“R1 matched o1”The R1 paper reports comparable results on selected reasoning tasks. That does not establish universal parity across tool use, safety, latency, long context, reliability or production operations.Overstated without task scope.
“R1 was fully open source”The release provided open weights, code and a report under permissive terms. The complete training data and reproducible end-to-end pipeline were not released.“Open-weight with code/report” is the precise description.
“DeepSeek succeeded without advanced Nvidia chips”The V3 report discloses a cluster of 2,048 Nvidia H800 GPUs and 2.788 million GPU-hours.False or materially overstated.
“DeepSeek had 50,000 H100s”A Thursday, January 23 Davos interview—four calendar days before the drop—aired the assertion without a public inventory or documentary proof.Unsupported public allegation.

Traditional media sometimes preserved scope—and sometimes amplified its loss

Nature's Thursday, January 23 feature—four calendar days before Nvidia's drop—kept the performance comparison tied to particular tasks and described the release as open-weight. TechCrunch's Sunday, January 26 roundup—one calendar day before the drop—reported the cost to train one model and embedded the social debate, yet its short comparison did not carry the V3 report's exclusion into the broader U.S. spending frame.

On Monday, January 27—the selloff day—CNBC's pre-market wording said a free model had been developed in two months for under $6 million. After the close that Monday, Fortune used similar R1-development language and linked to GitHub, although the repository does not support an all-in cost. Corrective analyst and investor commentary circulated on the same Monday; its timing matters because the simplified version had already become the dominant market shorthand.

Efficiency was still the real result

Correcting the accounting category does not erase DeepSeek's achievement. V3 and R1 demonstrated that algorithmic, architectural and training choices could move the cost-performance frontier under tighter hardware constraints. The investable question was never whether one bounded compute estimate equaled a total corporate budget. It was how lower unit costs might change total training and inference demand, competitive intensity, networking, power and the value of existing data-center plans.

This distinction is reusable: compare like with like, preserve the model name, preserve the training stage and identify whether a number is observed spend, estimated compute value, API price or a company-wide capital plan. The market-reaction study shows what happened when those categories collapsed into one Monday valuation thesis.

Did DeepSeek say R1 cost $6 million to build?

No. The $5.576 million estimate appeared in the V3 paper on Friday, December 27, 2024—thirty-one calendar days before Nvidia's drop—and covered the documented official V3 training work at an assumed $2 per GPU-hour. The R1 paper arrived on Wednesday, January 22, five calendar days before the drop, without an all-in development cost.

What did the $5.576 million estimate exclude?

The Friday, December 27, 2024 V3 report—thirty-one calendar days before Nvidia's drop—explicitly excluded earlier research and ablation experiments involving architecture, algorithms and data. It was not a total-company or total-program cost.

Was the low-cost result meaningless?

No. The documented efficiency was significant. The error was expanding a bounded V3 estimate into a claim about the full cost of R1 before Nvidia's selloff on Monday, January 27, 2025.

Sources

Never miss the next market-moving story

Seven market specialists, with experience dating back to 2006, watch global markets and U.S. stocks of every size. Start Pro to get the full live feed, clear context, measured price moves, search, watchlists, and alerts.

Start Pro — $29 for 7 days Watch live headlines -- free
See live news
The next market-moving story will not wait

See the important story while it still matters.

Seven market specialists bring experience dating back to 2006. MoveSurge adds the speed, coverage, and clear format built for today’s market.

Wide coverageGlobal markets and every size of U.S. stock Clear in secondsThe story, source, context, and measured move together Built on evidenceReal headlines, timestamps, prices, and trusted sources
Start Pro — $29 for 7 days View the live feed $29 today for 7 days. Then renews at the selected plan unless cancelled. Information only—no trade calls.