Contents
How much does an Nvidia H100 cost to buy? How much does it cost to rent an H100? What did Blackwell do to H100 prices? H100 vs H200 vs B200: which should you pay for? Buy, rent, or skip GPUs entirely? What is the real cost of an idle H100? Frequently asked questions about H100 pricing

Quick Answer

The Nvidia H100 price runs about $31,000 for a new 80GB card, $250,000 to $320,000 for a new 8-GPU HGX system, and meaningfully less refurbished. Renting costs $1.49 to $6.98 per GPU-hour depending on provider: Vast.ai marketplace hosts from $1.49, RunPod from $1.99, Lambda at $3.99, CoreWeave at $4.25, hyperscalers up to $6.98. Prices fell hard after Blackwell shipped. Verified August 20, 2026.

An H100 costs less today than at any point in its life, and it keeps getting cheaper. The B200 has shipped, AWS cut H100 rates roughly 44% in mid-2025 and the market followed, and the GPU that once required allocation lists and six-month waits is now the value play of AI infrastructure. Same silicon, completely different market, and most published prices haven’t caught up.

Which makes the H100 price question genuinely interesting, because there are now three right answers (buy new, buy used, rent) and the gap between the smart one and the default one is measured in hundreds of thousands of dollars.

Meanwhile GPU spend keeps climbing as the largest raw line in most AI budgets, and in CloudZero’s 2026 AI ROI survey of 260 finance leaders, 75% of those who cannot measure AI outcomes have held back investment, against 38% of those who can. GPU commitments are the biggest checks in AI. This page is the data.

How much does an Nvidia H100 cost to buy?

The numbers that matter, current market:

What you’re buyingPrice (August 2026)
Single H100 80GB card, new~$31,000 (live retail listings)
8x H100 HGX system, new$250,000 to $320,000 (~$285,000 typical)
8x H100 system, refurbishedSubstantially less: new runs about 70% more than equivalent used capacity
DGX H100 (Nvidia’s own 8-GPU system)Six figures, quote-based through partners

Prices this volatile are why year-stamped searches exist at all: the answer keeps changing. 2025’s number was roughly 44% higher before the cuts. Any price article without a visible date is quoting a different market.

Now the answer to the question behind the question: there is no official H100 MSRP. Nvidia has never published a list price; it sells through OEMs and system integrators, which is why every “H100 price” result on the internet is a third party publishing market observations, this page included. The difference is whether the observations are current and sourced.

The used market is the underrated story. Data center GPUs hold value far longer than the “obsolete next generation” narrative claims: Azure ran V100 instances for about 7.5 years, and CoreWeave reported H100 capacity coming off 2022-era contracts rebooking at 95% of original pricing (one operator’s data point, but a loud one). An H100 bought used today isn’t a depreciating brick; it’s discounted capacity with a long tail. Cheap rentals should have killed the case for owning GPUs. The residual-value data says otherwise.

How much does it cost to rent an H100?

The H100 GPU price per hour, on-demand, by provider:

ProviderPer GPU-hour (on-demand)Notes
Vast.ai (verified hosts)$1.49 to $2.27Marketplace; quality varies by host
RunPod$1.99 to $2.99Community vs secure cloud tiers
GCP (spot)~$2.25Interruptible; 60-91% off on-demand
Lambda$3.29 to $3.99SXM at the top of the range; no egress fees
CoreWeave$4.25 (PCIe) to $6.16 (per GPU on 8x HGX)InfiniBand, enterprise SLAs
AWS (p5.48xlarge, 8x H100)$55.04/hour per instance, ~$6.88 per GPUEffective rates lower with capacity blocks and commitments
AzureUp to $6.98Highest on-demand hyperscaler rate

Rates verified against provider pages and current market trackers as of August 20, 2026; spot capacity occasionally dips below $2 but is interruptible and not a number to budget on.

The pattern in that table has a name. Thunder Compute, itself a dedicated GPU cloud, puts median on-demand H100 pricing at about $4.17/hour on dedicated clouds versus $7.89 on hyperscalers, an 89% premium for the same silicon. You’re not paying more for the GPU; you’re paying for the enterprise wrapper (SLAs, networking, compliance, ecosystem), and whether that wrapper is worth 89% is a workload question, not a loyalty question.

Provider-by-provider rates all resolve to the table above, and all carry the same footnote: rates have moved repeatedly over the past eighteen months. Check the date on any table you trust, including this one.

Why the 5x spread for identical silicon? Four drivers. Form factor: SXM (faster interconnect, higher power) rents above PCIe on every provider, which is why RunPod and Lambda both show ranges.

  • Networking: InfiniBand-connected clusters for distributed training carry a premium single-GPU work never needs.
  • Egress: some providers charge nothing to move data out (Lambda) while hyperscalers meter it, a difference that can exceed the GPU rate on data-heavy jobs.
  • And commitment: reserved and spot capacity bracket every on-demand number in the table, from 60-91% spot discounts to reservation deals that cut on-demand by more than half. The on-demand price is the anchor, not the answer.

What did Blackwell do to H100 prices?

Repriced everything, in three moves:

Move one, the 2025 cuts. AWS cut H100 on-demand pricing 44% in June 2025, announced as up to 45% across its NVIDIA GPU instances, and the market followed. That wasn’t generosity; it was B-series inventory arriving and H100 demand curves bending.

Move two, the value repositioning. With B200 rentals running $3.99 to $5.50/hour on dedicated clouds (and ~$9.36 on AWS capacity blocks), the H100 stopped being the flagship and became the price-performance workhorse: mature software stack, abundant supply, falling rates.

Every framework, every tutorial, every debugging thread on the internet assumes an H100; that ecosystem maturity is a real discount that never appears on a rate card, because engineer-hours spent fighting drivers cost more than GPU-hours.

Move three, the math flip that surprises people. B200 delivers roughly 2.5x H100 training performance, so at $4.99/hour versus an H100 at $3.99, the newer chip is often cheaper per result despite the higher sticker. Hourly rate is the wrong metric; cost per training run is the right one, and it sometimes points at the more expensive GPU. This is the single most common error in GPU budgeting right now.

H100 vs H200 vs B200: which should you pay for?

GPUBuy (approx.)Rent (dedicated clouds)The honest positioning
H100 80GB~$31K/card; $250-320K per 8x system$1.99 to $4.25/hrThe value workhorse; best software maturity per dollar
H200Quote-based; premium over H100Provider-dependent, between H100 and B200More memory (141GB), same architecture; the inference-heavy pick
B200 (Blackwell)Reservation-heavy$3.99 to $5.50/hr (DataCrunch, Lambda, CoreWeave)~2.5x H100 training performance; often cheapest per result

The H200 and Blackwell pricing questions deserve their own breakdowns, and they are coming. The decision rule fits in a sentence: memory-bound inference favors H200, training throughput favors B200 per result, and everything cost-sensitive with a mature toolchain favors the H100 at 2026 rates.

The market’s dirty secret is that “which GPU” matters less than “how utilized,” and the next section is about exactly that.

Buy, rent, or skip GPUs entirely?

The three-way decision, with the math that decides it:

Rent when utilization is uncertain. At $3.99/hour, a rented H100 costs about $35,000 per GPU-year at full utilization, more than buying the card outright. But nobody runs full utilization from day one, and renting converts a six-figure guess into an hourly experiment you can end on Friday. The premium you pay per hour is the price of not being wrong at scale, which for a new workload is a good outcome on the market.

Buy (probably used) when utilization is proven. An 8x used system pays back against rentals in about seven months at full utilization: roughly $168,000 for used capacity against $23,300 a month to rent eight GPUs at $3.99 an hour around the clock. And the resale tail means the hardware isn’t worthless at the end. At 30% utilization, payback stretches to about two years, which is the polite way of saying: don’t buy for a workload you haven’t measured. And “measured” means measured: pull three months of actual GPU-hours from your rental invoices, not a capacity plan from a roadmap deck. The teams that regret buying didn’t get the hardware math wrong; they got their own utilization forecast wrong, usually by believing the optimistic version of it. Rental history is the only honest forecast you own.

Skip GPUs when an API does the job. A surprising money-saver in AI: many inference workloads running on rented H100s would cost a fraction as much on a per-token API, especially since OpenAI cut GPT-5.6 Luna by 80% and Terra by 20% on July 30, while GPU rates fell far less. The test is simple: price your monthly token volume against the LLM API pricing comparison and OpenAI’s rates, and only keep the GPUs if the infrastructure bill wins or the workload genuinely can’t leave your walls. Owning the stack is a strategy; defaulting to it is a tax, and in 2026 it’s a tax that got more expensive relative to the alternative.

Whichever route, the number that actually determines your H100 cost isn’t in any table above. It’s utilization.

What is the real cost of an idle H100?

An H100 at $3.99/hour costs the same whether it’s training a model or sitting idle while someone’s job queue is misconfigured. At 40% utilization, your effective H100 GPU cost isn’t $3.99 an hour; it’s nearly $10, which quietly puts your discount GPU cloud above hyperscaler list prices, meaning the procurement team won the negotiation and the idle time gave it all back. The market spent two years negotiating hourly rates down 44% while most teams’ idle time quietly ate a bigger discount than any procurement win.

And GPU spend has a visibility problem the rates hide. A team renting from Lambda for training, RunPod for experiments, and AWS for production inference gets three invoices in three units: Lambda bills per GPU-hour, AWS per instance-hour with eight GPUs bundled, RunPod per GPU-second. None of them can say what a single training run cost end to end, because the run touched all three. Add the inference bill from the models those GPUs produced, and the API spend for tasks that never justified self-hosting, and “what did our AI cost” becomes a reconstruction project.

Put numbers on it. A fine-tuning run using 8 H100s for 36 hours at Lambda’s $3.99: $1,149 per run. Same run on an idle-prone shared cluster where jobs queue and GPUs wait: the wall-clock stretches to 60 hours and the run costs $1,915, a 67% premium with no invoice line explaining it. Now multiply by the runs nobody tracked because each one looked small. GPU waste doesn’t arrive as a big number; it arrives as a thousand reasonable-looking ones.

That’s the problem CloudZero works on, and the mechanics are specific to GPUs.

Native AWS, Azure, and GCP integrations pull hyperscaler GPU spend, and the AnyCost framework ingests specialized GPU cloud bills into the same normalized view in the AI Hub.

Dimensions allocate that spend to cost per training run, per model, per team, and per experiment, no tags required, which converts a single monthly GPU number into a per-run one: what each model costs per training run, and whether the run count moved.

Budgets and forecasting put forward numbers on the biggest checks in AI, and anomaly detection with hour-level data catches the misconfigured job that left 64 GPUs running over a weekend, before the invoice announces it.

The industry got very good at negotiating GPU hourly rates and stayed very bad at knowing what the hours produced. Only one of those skills shows up in AI ROI. The full pattern, with spend data across the market, is in CloudZero’s State of AI Costs report and the AI spend management guide.

Schedule a demo to see GPU spend allocated to cost per training run and cost per model, get a free cloud cost assessment to find your effective per-GPU-hour rate including the idle time, or take a self-guided tour to explore AI cost allocation in action. More customer stories on the customers page.

Frequently asked questions about H100 pricing