AI debate · Reviewed September 18, 2026

What a Frontier Model Costs to Train, and Who Profits From Slowing Down

By AI Agent Hub Editorial Desk · Review method · Corrections

Scope: every training cost figure on this page is a third-party estimate, not a lab disclosure. No frontier lab publishes its actual training spend. Where sources disagree, the disagreement is shown rather than averaged away.

The September 2026 argument over whether AI development should slow is usually presented as a disagreement about risk. It is also a disagreement about money, and reading it that way explains the alliance pattern that otherwise looks strange: the companies that pay for training want to slow down, and the companies that sell the computing do not.

The cost curve

These figures are compiled from public estimates, hardware analysis, and the small number of partial disclosures. They are directionally useful and individually soft.

ModelYearEstimated training costNote
GPT-3 (175B)2020~$4.6 millionEstimate
PaLM (540B)2022~$12 millionEstimate
GPT-42023~$79 millionEstimate; other sources cite $78–100 million
Gemini Ultra2024~$191 millionEstimate
Llama 3.1 405B2024~$60 million or ~$170 millionSources disagree by nearly 3x; treat as unresolved
DeepSeek V32024–2025~$5.6 millionEstimate
DeepSeek R1 post-training2025~$294,000Excludes base model cost; see below
Frontier tier2026$500 million – $1.5 billionEstimate range; not disclosed

The Llama row deserves attention because it is not a rounding difference. One widely circulated table puts Llama 3.1 405B at roughly $60 million, another at roughly $170 million. Both are presented as estimates for the same model. That gap is a reminder of what these numbers are: reconstructions from cluster size, duration, utilization assumptions, and hardware pricing, not accounts.

Where the money goes

Training spend is not one line item. Published breakdowns converge on a similar split:

One consequence is that "training cost" is not a single well-defined quantity. A figure that includes only rented GPU hours and a figure that includes staff and data can differ by a third for the same model, which is part of why published estimates diverge.

The DeepSeek counterexample, and why it does not mean what it sounds like

DeepSeek R1 is routinely cited as having been trained for about $294,000. Against a $500 million frontier budget, that is a ratio of roughly 1,700 to 1, and it is used to argue that spending is not buying capability.

The comparison needs a correction before it can be used. The $294,000 figure describes a post-training run performed on top of an existing base model. It does not include the cost of producing that base model, which is the expensive part. Comparing it to a full pre-training budget is comparing a renovation to a construction.

The weaker version of the claim still holds, and it matters more: efficiency choices change the cost of capability by large multiples. DeepSeek V3 at roughly $5.6 million is a full training run, and while it is not frontier-tier on every axis, the gap between $5.6 million and $500 million is not explained by a proportional gap in usefulness. That has a direct implication for the slowdown debate. If capability were strictly a function of spend, then limiting spend would reliably limit capability, and a speed limit would be a clean instrument. It is not, so a spending cap is a blunt instrument at best.

Who gains from a slower frontier

Follow the revenue, and the September 2026 alignment stops looking puzzling.

Position in the stackRevenue depends onEffect of a slower frontier
Chip and hardware suppliersVolume of accelerators soldReduced growth in demand; this is why the Nvidia CEO rejects the premise
Hyperscale cloud providersCompute rented, including for trainingLower training demand, though inference demand is less affected
Frontier model labsModel capability and adoptionLower cost and lower risk exposure, at the price of slower differentiation
API consumers and developersToken price and reliabilityLargely unaffected short term; affected if consolidation reduces competition on price

Note the asymmetry. A lab that slows down gives up a competitive race but keeps its product and cuts its largest cost line. A hardware supplier that slows down loses revenue growth and gets nothing back. Both positions can be held sincerely; that does not make them equally costly to hold.

The market read this immediately. On September 14, 2026, two days after the pacing essay appeared, AI-related equities fell sharply across Asian markets, with SoftBank reported down as much as 13.2% and semiconductor names including SK Hynix, Samsung Electronics, and Tokyo Electron also declining. Investors did not need to form a view on AI risk to notice that a slower frontier changes the value of infrastructure built for a faster one.

What this means if you buy tokens instead of building models

For most readers of this site, training cost is not a line item — token price is. The two are connected, but weakly and slowly:

Estimate any of this directly with the AI API cost calculator, or check how much text a given budget buys using the token counter. Both use published rates and show their assumptions.

Frequently asked questions

How much does it cost to train a frontier AI model in 2026?

Public estimates for a 2026 frontier-tier run range from about $500 million to $1.5 billion. No lab publishes actual training costs, so these are third-party reconstructions rather than audited figures.

Why do estimates differ so much between sources?

Because they depend on assumptions about cluster size, training duration, hardware utilization, whether hardware is owned and amortized or rented at cloud rates, and whether staff and data costs are included. Published figures for the same model can differ by a factor of two or more.

What share of training cost is GPU compute?

Published breakdowns put compute at roughly 60 to 70 percent, with data preparation around 15 percent, engineering around 12 percent, and infrastructure around 8 percent.

Did DeepSeek really train a competitive model for under $300,000?

The widely cited $294,000 figure for DeepSeek R1 covers a post-training run on an existing base model, not the base model itself. It should not be compared directly with a full frontier pre-training budget.

Why do chip companies and model labs disagree about slowing AI down?

Their revenue is tied to different quantities. Hardware and cloud suppliers earn more as compute demand grows, while frontier labs carry the training bill and the reputational risk. A slower frontier reduces revenue growth for one and cost and risk exposure for the other.

Primary sources

Bottom line

Training costs are large, soft, and disputed — the same model is credibly estimated at $60 million and $170 million depending on the source. What is not disputed is the incentive structure: the companies paying for frontier training and the companies selling the computing want opposite things from a slowdown, which is why the September 2026 debate split along supply-chain lines rather than ideological ones. For anyone buying tokens rather than building models, the practical variable remains published unit rates and how much competition keeps them down.