What a Frontier Model Costs to Train, and Who Profits From Slowing Down
Scope: every training cost figure on this page is a third-party estimate, not a lab disclosure. No frontier lab publishes its actual training spend. Where sources disagree, the disagreement is shown rather than averaged away.
The September 2026 argument over whether AI development should slow is usually presented as a disagreement about risk. It is also a disagreement about money, and reading it that way explains the alliance pattern that otherwise looks strange: the companies that pay for training want to slow down, and the companies that sell the computing do not.
The cost curve
These figures are compiled from public estimates, hardware analysis, and the small number of partial disclosures. They are directionally useful and individually soft.
| Model | Year | Estimated training cost | Note |
|---|---|---|---|
| GPT-3 (175B) | 2020 | ~$4.6 million | Estimate |
| PaLM (540B) | 2022 | ~$12 million | Estimate |
| GPT-4 | 2023 | ~$79 million | Estimate; other sources cite $78–100 million |
| Gemini Ultra | 2024 | ~$191 million | Estimate |
| Llama 3.1 405B | 2024 | ~$60 million or ~$170 million | Sources disagree by nearly 3x; treat as unresolved |
| DeepSeek V3 | 2024–2025 | ~$5.6 million | Estimate |
| DeepSeek R1 post-training | 2025 | ~$294,000 | Excludes base model cost; see below |
| Frontier tier | 2026 | $500 million – $1.5 billion | Estimate range; not disclosed |
The Llama row deserves attention because it is not a rounding difference. One widely circulated table puts Llama 3.1 405B at roughly $60 million, another at roughly $170 million. Both are presented as estimates for the same model. That gap is a reminder of what these numbers are: reconstructions from cluster size, duration, utilization assumptions, and hardware pricing, not accounts.
Where the money goes
Training spend is not one line item. Published breakdowns converge on a similar split:
- Compute, roughly 60–70%. GPU or TPU time, owned and amortized or rented. The single most provider-sensitive component.
- Data preparation, roughly 15%. Crawling, filtering, deduplication, tokenization, and human labeling for preference training.
- Engineering, roughly 12%. The research and infrastructure staff who design, run, and debug a months-long training run.
- Infrastructure, roughly 8%. Storage, high-speed networking, orchestration, and monitoring.
One consequence is that "training cost" is not a single well-defined quantity. A figure that includes only rented GPU hours and a figure that includes staff and data can differ by a third for the same model, which is part of why published estimates diverge.
The DeepSeek counterexample, and why it does not mean what it sounds like
DeepSeek R1 is routinely cited as having been trained for about $294,000. Against a $500 million frontier budget, that is a ratio of roughly 1,700 to 1, and it is used to argue that spending is not buying capability.
The comparison needs a correction before it can be used. The $294,000 figure describes a post-training run performed on top of an existing base model. It does not include the cost of producing that base model, which is the expensive part. Comparing it to a full pre-training budget is comparing a renovation to a construction.
The weaker version of the claim still holds, and it matters more: efficiency choices change the cost of capability by large multiples. DeepSeek V3 at roughly $5.6 million is a full training run, and while it is not frontier-tier on every axis, the gap between $5.6 million and $500 million is not explained by a proportional gap in usefulness. That has a direct implication for the slowdown debate. If capability were strictly a function of spend, then limiting spend would reliably limit capability, and a speed limit would be a clean instrument. It is not, so a spending cap is a blunt instrument at best.
Who gains from a slower frontier
Follow the revenue, and the September 2026 alignment stops looking puzzling.
| Position in the stack | Revenue depends on | Effect of a slower frontier |
|---|---|---|
| Chip and hardware suppliers | Volume of accelerators sold | Reduced growth in demand; this is why the Nvidia CEO rejects the premise |
| Hyperscale cloud providers | Compute rented, including for training | Lower training demand, though inference demand is less affected |
| Frontier model labs | Model capability and adoption | Lower cost and lower risk exposure, at the price of slower differentiation |
| API consumers and developers | Token price and reliability | Largely unaffected short term; affected if consolidation reduces competition on price |
Note the asymmetry. A lab that slows down gives up a competitive race but keeps its product and cuts its largest cost line. A hardware supplier that slows down loses revenue growth and gets nothing back. Both positions can be held sincerely; that does not make them equally costly to hold.
The market read this immediately. On September 14, 2026, two days after the pacing essay appeared, AI-related equities fell sharply across Asian markets, with SoftBank reported down as much as 13.2% and semiconductor names including SK Hynix, Samsung Electronics, and Tokyo Electron also declining. Investors did not need to form a view on AI risk to notice that a slower frontier changes the value of infrastructure built for a faster one.
What this means if you buy tokens instead of building models
For most readers of this site, training cost is not a line item — token price is. The two are connected, but weakly and slowly:
- Your bill follows published unit rates, not anyone's training budget. A $500 million training run does not entitle a lab to charge more per token; prices are set by competition among providers.
- Consolidation is the channel that matters. If only a few organizations can fund frontier training, competition on price weakens over time. That is the mechanism by which training economics eventually reaches your invoice.
- Efficiency gains reach you faster than training costs do. Falling inference cost per token has been the dominant trend in published rates, visible in the spread between DeepSeek V4.1 Flash at $0.15 per million input tokens and Claude Opus 5 at $5.00.
Estimate any of this directly with the AI API cost calculator, or check how much text a given budget buys using the token counter. Both use published rates and show their assumptions.
Frequently asked questions
How much does it cost to train a frontier AI model in 2026?
Public estimates for a 2026 frontier-tier run range from about $500 million to $1.5 billion. No lab publishes actual training costs, so these are third-party reconstructions rather than audited figures.
Why do estimates differ so much between sources?
Because they depend on assumptions about cluster size, training duration, hardware utilization, whether hardware is owned and amortized or rented at cloud rates, and whether staff and data costs are included. Published figures for the same model can differ by a factor of two or more.
What share of training cost is GPU compute?
Published breakdowns put compute at roughly 60 to 70 percent, with data preparation around 15 percent, engineering around 12 percent, and infrastructure around 8 percent.
Did DeepSeek really train a competitive model for under $300,000?
The widely cited $294,000 figure for DeepSeek R1 covers a post-training run on an existing base model, not the base model itself. It should not be compared directly with a full frontier pre-training budget.
Why do chip companies and model labs disagree about slowing AI down?
Their revenue is tied to different quantities. Hardware and cloud suppliers earn more as compute demand grows, while frontier labs carry the training bill and the reputational risk. A slower frontier reduces revenue growth for one and cost and risk exposure for the other.
Primary sources
- Epoch AI — training compute and cost analysis, the most frequently cited primary source for this data
- GPUnex — AI training costs 2026 — cost timeline and breakdown percentages
- Capital & Compute — what it costs to train AI models (2026) — frontier-tier ranges, explicitly labeled as estimates
- AI Agent Hub verified catalog — published API rates used for the token conversions on this page
- Pacing the Frontier debate — the September 2026 policy argument this economics sits behind
Bottom line
Training costs are large, soft, and disputed — the same model is credibly estimated at $60 million and $170 million depending on the source. What is not disputed is the incentive structure: the companies paying for frontier training and the companies selling the computing want opposite things from a slowdown, which is why the September 2026 debate split along supply-chain lines rather than ideological ones. For anyone buying tokens rather than building models, the practical variable remains published unit rates and how much competition keeps them down.