The Real Cost of Training an AI Model, Broken Down

GPUs, electricity, and data — what it actually takes to build a frontier model, and why the bill keeps climbing

Last updated: July 2026

Every time a new frontier model ships — GPT-4, Gemini Ultra, Claude, Llama, DeepSeek — a number gets attached to it in the press: "$100 million," "$5 million," "over a billion by 2027." These numbers are usually right and misleading at the same time, because "training cost" means different things depending on who's doing the counting. A GPU-hour bill is not the same as a total R&D budget, and a press release is not an audited financial statement.

This piece breaks the real cost structure into its actual components — silicon, electricity, data, and people — using the most rigorous public research available (mainly Epoch AI's cost model and the peer-reviewed paper behind it), plus current 2026 hardware and cloud pricing. Along the way I'll flag where the popular numbers ("GPT-4 cost $100M," "DeepSeek cost $5.5M") are technically true but structurally incomplete.


1. The four-part cost stack

The best public accounting of frontier training costs comes from Epoch AI, a research institute that built a detailed cost model covering 45 frontier systems released since 2016. Their arXiv paper, "The Rising Costs of Training Frontier AI Models", breaks a training run into these buckets:

  • AI accelerator chips (GPUs/TPUs) — the single largest line item, roughly 47–67% of total cost
  • Servers and cluster-level networking/interconnect — another 15–22% for servers, 9–13% for interconnect
  • Energy — a surprisingly small 2–6% of the amortized total
  • R&D staff — 29–49% of cost, depending on the lab and how compute-efficient their hardware is

Their headline finding: amortized training costs for the most compute-intensive models have grown by roughly 2.4x per year since 2016, and if that trend continues, the largest training runs will cost more than a billion dollars by 2027. A separate cloud-rental-based method, which is simpler but tends to overstate true costs because most labs use owned rather than rented hardware, found a similar 2.6x annual growth rate, with cost estimates roughly twice as high on average as the hardware-and-energy method.

Now let's take each pillar apart.


2. GPUs: the line item everyone talks about

What the chips actually cost

Nvidia doesn't publish list prices for its data-center GPUs, so figures come from resellers, system integrators, and cloud providers. As of mid-2026:

  • An H100 80GB (the chip that trained most current frontier models, including Llama 3.1) runs roughly $25,000–$40,000 per unit depending on form factor (PCIe vs. SXM5), per OpsLyft's 2026 pricing breakdown.
  • A full DGX H100 8-GPU server system costs $350,000–$400,000+, according to the same source and corroborated by CloudZero's H100 cost analysis.
  • The current-generation Blackwell B200 sells for a similar $30,000–$40,000 per card, but Epoch AI's December 2025 hardware-cost analysis estimated its actual production cost at only $5,700–$7,300 — implying Nvidia is running close to an 80% gross margin on the chip, as reported by this Medium analysis of AI compute economics.

Renting instead of buying

Most labs — especially smaller ones — rent GPU-hours from cloud providers rather than buying hardware outright. Pricing varies enormously by vendor:

Provider tier H100 on-demand ($/GPU-hr) Notes
Hyperscalers (AWS, Azure, GCP) ~$3.00–$12.29 AWS only sells H100 in 8-GPU instances, so small jobs overpay
Neo-clouds (Lambda, CoreWeave, RunPod) ~$2.25–$3.44 Purpose-built for AI, 50–75% cheaper than hyperscalers
Spot/marketplace (Vast.ai, Spheron) ~$0.34–$2.00 Cheapest, but capacity can be reclaimed with little notice

Source figures compiled from Thunder Compute's July 2026 pricing tracker, GetDeploying's live H100 comparison, and CloudZero's cross-cloud comparison. Notably, AWS cut H100 rental pricing by roughly 44% in a single month in June 2025 as Blackwell hardware became available, which is a good reminder that any GPU-hour price quoted today will likely be stale within a year.

What a real cluster costs

Do the arithmetic on a serious training run and the numbers get large fast. Renting a 25,000-GPU H100 cluster for 90 continuous days — a realistic scale for a 2024–2025-era frontier model — runs to roughly $160 million in pure compute rental, even before staff, data, or failed experiments are counted (estimate via this infrastructure cost analysis, consistent with Epoch AI's underlying methodology).

At the frontier's bleeding edge, the clusters are bigger still. xAI's Colossus facility in Memphis scaled from 100,000 to roughly 200,000 H100-equivalent chips through 2025. Meta's internal H100-class capacity exceeded 350,000 H100-equivalents by the end of 2025. These are not rented — they're owned, which shifts the accounting from an hourly rental rate to depreciation, power contracts, and real estate, described well in this frontier-training-cost overview.

My take: the GPU story is really two separate stories layered on top of each other. One is Moore's-Law-style progress — compute-per-dollar keeps improving, and Epoch AI's own trend data shows performance-per-dollar rising at roughly 37% per year. The other is a demand shock: everyone wants the same chips at the same time, so scarcity premiums, margin capture by Nvidia, and bidding wars between hyperscalers push prices in the opposite direction. Which force wins depends on your time horizon — near-term, scarcity dominates; over 3-5 years, efficiency usually wins.


3. Electricity: real, rising, but not the biggest line item

Energy gets outsized media attention relative to its actual share of training cost — Epoch AI puts it at just 2–6% of the amortized total. But the absolute numbers are still large enough to matter for grid planning and climate accounting.

How much power a training run actually uses

Public disclosures are rare, so most figures are estimates:

  • GPT-3 (2020): widely cited at around 1,287 megawatt-hours (MWh), based on a Google/UC Berkeley study, with roughly 502–552 metric tons of CO2e emitted — equivalent to more than 200 round trips between Paris and New York by plane, according to IFP School's energy explainer.
  • GPT-4-class models: estimates diverge sharply depending on methodology. Independent analysis and commentary (including Scott Alexander's widely-read piece on scaling economics) puts GPT-4's training energy use at roughly 50 gigawatt-hours (GWh). Epoch AI's own more recent estimate for GPT-4o-class models found training runs consumed around 20–25 megawatts of continuous power over about three months — which works out to roughly 43,000–54,000 MWh, broadly consistent with the ~50 GWh figure. (Lower single-digit-thousand-MWh estimates circulating elsewhere appear to use different assumptions about GPU count and utilization, so treat any single figure with some skepticism.)
  • For comparison, that's enough electricity to power roughly 20,000–40,000 average American homes for a year — for one training run.

The bigger picture: data centers, not just training

Training a single model is a one-time (if enormous) energy draw. The ongoing story is inference — running the model for billions of daily queries — which now dwarfs training energy for any popular product. Epoch AI's own analysis, after correcting for outdated assumptions, found a typical GPT-4o query consumes roughly 0.3 watt-hours, about ten times less than an earlier viral estimate — but at OpenAI's query volume, that still adds up to hundreds of megawatts of continuous inference load.

At the macro level: data centers currently account for roughly 1–2% of global electricity consumption (300–400 TWh/year), and the IEA projects this could double by 2030, driven overwhelmingly by AI. This is why Microsoft signed a 20-year deal to restart Three Mile Island's Unit 1 reactor, and why Google, Amazon, and others are now signing nuclear and small-modular-reactor power purchase agreements specifically for AI campuses.

My take: the "AI is destroying the planet" framing and the "electricity is a rounding error" framing are both technically defensible and both miss the point. Energy is a small percentage of training cost, but it is the hard physical constraint on how fast the industry can scale — you can raise capital and buy chips faster than you can permit and build a gigawatt of new grid capacity. Multiple 2026 industry analyses (e.g., Epoch AI's trends dashboard) now describe power, not chip supply, as the binding constraint on further scaling. That's a more interesting and more important story than the per-query watt-hour debates that dominate headlines.


4. Data: the least visible, fastest-growing cost

Five years ago, training data was essentially free — scrape the public web, done. That era is over, for two overlapping reasons: legal risk (a wave of copyright lawsuits, including the New York Times v. OpenAI/Microsoft suit) and simple scarcity (frontier models now consume text faster than the internet produces new material).

What data licensing actually costs

A real market has emerged, with wildly different price points depending on the content and buyer:

  • Shutterstock earned an estimated $104M in 2023, rising to $138M in 2024 in AI licensing revenue, with individual Big Tech deals in the $25M–$50M range each (per Bloomberg, via Quartz's pricing breakdown).
  • Reddit's deal with Google is reportedly worth about $60M/year; its OpenAI deal is estimated around $70M/year, per Columbia Journalism Review's reporting.
  • News Corp's OpenAI deal reportedly averages about $50M/year across its portfolio of the Wall Street Journal, New York Post, and other titles; Amazon's deal with the New York Times is valued at $20M–$25M/year.
  • Academic publisher Wiley licensed its back catalog of papers for a one-time $23M.
  • The overall AI training-dataset licensing market was valued at roughly $3.6–4.8 billion in 2025, and one estimate projects it could reach $22.6 billion by 2034 (MarketIntelo).

Deal structures are also shifting. Early 2023–2024 agreements were mostly flat annual training-rights fees; by 2025–2026 the market moved toward usage-based and "grounding" (real-time retrieval) pricing, since most frontier labs consider their models already reasonably well-trained on public text and now want ongoing access to fresh content rather than one-time bulk archives — a trend documented in detail by Media & the Machine's licensing tracker.

Licensing fees are the visible, negotiated cost of data. There's a second, much larger and much less predictable cost: litigation exposure for training on unlicensed material. This is a real budget line now, not a hypothetical one — courts are actively deciding whether large-scale scraping for AI training qualifies as fair use, and settlements can dwarf any licensing deal on the table.

My take: data is the pillar most likely to reshape the industry's structure over the next five years. Compute costs scale down over time (Moore's Law, better chips, more efficient architectures). Data scarcity does the opposite — the stock of unique, high-quality, legally clean human text is fixed, and every lab is drawing from the same shrinking pool simultaneously. That's a genuinely different kind of bottleneck than "buy more GPUs," and it's part of why synthetic data generation (using a frontier model to generate training signal for its successor) has become such a central technique rather than a stopgap.


5. People: the quiet 30–50%

R&D staff costs are the least-discussed pillar but frequently rival hardware costs in scale. Epoch AI's paper found that for GPT-4, the number of disclosed contributors rose to 284 people, and amortized hardware cost over the whole development process reached about $90 million — with staff costs adding substantially on top. For Gemini Ultra, R&D staff costs made up the highest share Epoch AI measured — an estimated 49% of total cost — partly because Google's TPUs are cheaper for Google than commercial GPU pricing, which shrinks the hardware side of the ratio.

This matters because elite AI research talent is genuinely scarce and genuinely expensive. Compensation packages for top researchers at frontier labs can run into the millions of dollars annually in total comp (salary plus equity), and DeepSeek — often cited as a "cheap" lab — reportedly offers salaries north of $1.3 million to recruit competitively against Chinese Big Tech.

My take: this is the cost line that most explains why "just rent some GPUs" doesn't get you a frontier model. You can rent the same H100 cluster DeepSeek used for a few million dollars. You cannot rent the years of accumulated systems-engineering knowledge — efficient distributed training frameworks, data curation pipelines, evaluation infrastructure — that make that cluster produce a competitive model instead of a mediocre one.


6. Case study: what "DeepSeek cost $5.6 million" actually means

DeepSeek's V3 technical report became one of the most-cited (and most-misunderstood) cost claims in AI. It's worth unpacking carefully because it's a perfect illustration of how "training cost" headlines can be true and incomplete at once.

What DeepSeek actually disclosed: V3 was trained on 2,048 Nvidia H800 GPUs for about two months, totaling 2.788 million GPU-hours. At an assumed rental rate of $2/GPU-hour, that's $5.576 million — a figure DeepSeek itself explicitly labeled as covering only the final training run, excluding "costs associated with prior research and ablation experiments on architectures, algorithms, or data".

What that number leaves out, according to independent analysis by SemiAnalysis and The Register:

  • The actual hardware purchase cost of DeepSeek's GPU fleet — estimated at $500M+ in total server capex, including roughly 10,000 A100s and 50,000 H800/H100-class chips
  • R&D staff costs for DeepSeek's technical team, which likely run into the tens of millions of dollars annually given the reported $1.3M+ compensation for top researchers
  • All the failed runs, architecture experiments, and data pipeline work that happened before the final successful training run
  • The R1 reasoning model that built on top of V3 (some outlets confused a much smaller reinforcement-learning-phase compute figure — around $294,000 — with the full training cost, when in fact it excluded the 2.79 million GPU-hours needed to build the V3 base model underneath it)

A widely cited independent estimate puts DeepSeek's real all-in annual operating cost closer to $500 million–$1.3 billion, not $5.6 million — not because DeepSeek lied, but because the $5.6M figure was never meant to represent total cost in the first place.

My take: this is the single most important thing to understand about every training-cost headline you'll ever read, including the ones in this article. "Training cost" almost always means the marginal compute cost of the final, successful run — not the total cost of the research program that made that run possible. When Sam Altman says GPT-4 "cost more than $100 million," he is very likely also only describing this narrower slice. The real all-in cost of running a frontier lab — salaries, failed experiments, infrastructure, data, legal exposure — is a multiple of the number that makes headlines, often 2-5x by informed industry estimates.


7. What frontier models have actually cost — a comparison table

Figures below combine Epoch AI's cost-model estimates (via the Stanford AI Index and Epoch AI's own dataset), company disclosures, and independent reporting. Treat every number as an estimate with real uncertainty bands, not an audited figure — none of these labs publish full financials for a specific model.

Model Year Estimated training-run cost Notes
Transformer (original) 2017 ~$670 Baseline for scale comparison
GPT-3 2020 ~$2–4 million 175B parameters
PaLM 2022 ~$3–12 million Compute only
OPT-175B 2022 ~$1.5–2 million One of the few models with disclosed cluster logbook costs
GPT-4 2023 ~$78–100+ million Sam Altman confirmed "more than $100M"; Stanford AI Index compute-only estimate is $78M
Gemini Ultra 1.0 2023 ~$30–192 million (hardware/energy only); ~$191M cited by Stanford AI Index R&D staff share reportedly the highest of any model measured (49%)
Claude 3.5 Sonnet 2024 "a few tens of millions" Per Anthropic CEO Dario Amodei's own public statement
Llama 3.1 405B 2024 Tens of millions (compute); trained on H100s Meta has not disclosed a specific dollar figure
DeepSeek V3 Dec. 2024 $5.6M (final-run compute only); likely $500M+ all-in See case study above
GPT-5 / Gemini-Ultra-class 2026 models 2025–2026 Estimated $200–500 million Industry estimates, not confirmed by labs
Next frontier tier (projected) 2027 $1–3 billion Epoch AI trend projection if 2.4x/year growth continues

Sources: Epoch AI cost paper, Forbes/Statista coverage of Epoch AI's 2024 release, Galileo's LLM cost analysis, and the DeepSeek sources cited in Section 6.


8. The infrastructure arms race behind the models

None of the per-model numbers above capture the real scale of capital being deployed. The four largest US hyperscalers — Amazon, Google, Meta, and Microsoft — are projected to spend a combined ~$725 billion on capital expenditure in 2026, up roughly 77% from ~$410 billion in 2025, according to multiple trackers compiling company earnings guidance (ValueAddVC's live capex dashboard; figures corroborated by Futurum Group). Individually:

  • Amazon: ~$200 billion
  • Google/Alphabet: ~$175–185 billion
  • Meta: ~$115–135 billion
  • Microsoft: ~$110–120 billion

Add Oracle's Stargate commitments (a $500 billion joint venture with OpenAI and SoftBank, deployed over several years) and total US AI infrastructure commitments for 2026 alone approach $700 billion. Goldman Sachs now projects a combined $5.3 trillion in hyperscaler capex from 2025 through 2030.

This is the context that makes any single model's training cost look almost modest by comparison — a $500 million training run is a rounding error against a $725 billion annual infrastructure budget, because that budget also has to fund years of inference serving, R&D experimentation, and the next three generations of hardware.

My take: this is the number that should actually worry people more than any individual model's price tag. A $500M training run is a large but bounded, one-time expense. A $700B/year ongoing capital commitment, funded partly through debt (hyperscalers reportedly raised over $100 billion in debt in 2025 alone for this buildout), is a bet that AI revenue will eventually justify infrastructure spending at a scale roughly comparable to entire national R&D budgets. If that bet is wrong, the fallout won't be contained to one company's balance sheet.


9. So, what does it actually cost to train a frontier model?

Pulling it together, a realistic 2026-era frontier model likely breaks down something like this:

  1. Final training run (compute + energy): $50–500 million, depending on model scale
  2. Failed runs, ablations, and scaling experiments before the final run: often 10–30% of the FLOPs of the successful run, adding meaningfully to the true program cost
  3. R&D staff (including equity comp) for the team and the months/years of work that produced the final architecture: commonly 30–50% on top of hardware costs
  4. Data acquisition and licensing: anywhere from effectively $0 (public web scraping, accepting legal risk) to tens of millions of dollars in negotiated licensing deals
  5. Legal/compliance exposure: unpredictable, but potentially enormous — settlements in copyright litigation have already reached into the billions for at least one major lab

Add it up, and the "headline" training cost that makes news is routinely a fraction — often less than half — of the true all-in cost of getting a frontier model built.


Final take

The industry narrative tends to pick whichever number serves the story: labs highlight low compute-only figures to seem efficient (DeepSeek), or high headline figures to justify enormous funding rounds (frontier labs raising billions). Both can be true simultaneously, because they're measuring different things. If you remember one framework from this article, make it this: ask what's excluded, not just what's included, whenever you see a training-cost number in a headline.

The trajectory is unambiguous regardless of which exact number you trust: costs are compounding at roughly 2.4–3x per year, the industry itself is now betting over half a trillion dollars annually that this is worth it, and the binding constraint is shifting away from raw silicon availability and toward electricity, high-quality data, and skilled people — three things money can't instantly buy more of.


Sources referenced in this piece

Note: figures throughout this piece are estimates from independent research and public disclosures. AI labs generally do not publish audited, itemized training costs, so treat specific dollar figures as informed approximations rather than confirmed accounting.