Published: October 2026 · Pricing last verified: October 10, 2026. Provider pricing and availability may change; check the linked source before committing to a workload.
GPU cloud pricing for the same class of NVIDIA H100 runs from $2.89 per GPU-hour on RunPod (a single H100 PCIe pod) to about $11.68 per GPU-hour on Google Cloud, where an eight-GPU a3-megagpu-8g VM lists at $93.40 an hour. Prices were checked on October 10, 2026. That is a 4x spread for one chip, but the comparison is imperfect: the cheap end is one PCIe card with 16 vCPUs, the expensive end is a full SXM node with 3,200 Gbps of networking. Newer chips sit higher still, with NVIDIA B200 at $6.69 to $14.21 per GPU-hour depending on provider and billing model.
The more useful finding is that the hourly rate is the wrong number to optimise. On a published benchmark, the most expensive chip per hour produced the cheapest tokens.
Key Takeaways
- On-demand H100 pricing spans $2.89 to $11.68 per GPU-hour across the providers checked; H200 runs $5.29 to $10.60 and B200 $6.69 to $8.60 where an on-demand rate is published.
- Google Cloud now lists Ironwood (TPU7x) at $12.00 per chip-hour on demand in us-central1, falling to $5.40 on a three-year commitment. TPU chip-hours are not equivalent to GPU-hours.
- Spot and commitment discounts are larger than most provider-to-provider gaps. CoreWeave’s spot H100 node costs 60% less than its on-demand rate.
- Using MLPerf Llama 2 70B results, an eight-GPU B200 node produced output tokens for roughly $0.15 to $0.31 per million, against $0.41 to $0.64 for an H100 node. Throughput, not rental rate, set the winner.
- Utilisation is the cost lever buyers underestimate. At 60% productive use, a $0.48 compute cost per million tokens becomes $0.83 once modest overhead is added.
- Price your workload in cost per million output tokens at your realistic utilisation before comparing any two quotes.
Quick Navigation
- The GPU cloud pricing comparison: H100, H200, B200 and TPU7x
- What the hourly GPU cloud price does not tell you
- On-demand, reserved and spot GPU cloud pricing
- From GPU cloud pricing to cost per million tokens
- A reusable GPU cost calculator
- Which provider fits which workload
- How to check a GPU cloud quote before you sign
- Conclusion
- Frequently Asked Questions
- How much does it cost to rent an H100 GPU per hour in 2026?
- Is an H200 more expensive than an H100?
- How much does an NVIDIA B200 cost per hour?
- Does Google Cloud publish TPU v7 pricing?
- Is spot GPU pricing worth the interruption risk?
- Is renting a GPU cheaper than using an AI API?
- How do I calculate GPU inference cost per million tokens?
- What additional costs should I include in a GPU cloud budget?
The GPU cloud pricing comparison: H100, H200, B200 and TPU7x
Every figure below comes from the provider’s own pricing page, checked October 10, 2026. Where a provider sells only eight-GPU nodes, the per-GPU figure is the node price divided by eight, shown next to the original. US regions are used unless stated. Azure and AWS on-demand rates are the exception: both companies render prices in an interactive calculator, so those two rows rely on third-party mirrors of the official price APIs and are marked as such.
Table: scroll sideways to view all columns when needed.
| Provider | Accelerator | Configuration | On-demand | Reserved / committed | Spot / interruptible | Region and conditions | Verified | Source |
|---|---|---|---|---|---|---|---|---|
| RunPod | H100 PCIe | 1 GPU pod, 80 GB, 16 vCPU, 188 GB RAM | $2.89/GPU-hr | Quote required | Not verified | Per-second billing; Community vs Secure tier not distinguished in page text | 2026-10-10 | RunPod pricing |
| RunPod | H100 SXM | 1 GPU pod, 20 vCPU, 125 GB RAM | $3.99/GPU-hr | Quote required | Not verified | As above | 2026-10-10 | RunPod pricing |
| Lambda | H100 SXM | 8x instance, 208 vCPU, 1,800 GiB RAM | $3.99/GPU-hr ($31.92/node) | Quote required | Not offered | Self-serve, first come; plus sales tax | 2026-10-10 | Lambda pricing |
| AWS | H100 | p5.48xlarge, 8 GPU | $55.04/node ($6.88/GPU), via third-party mirror | Capacity Block $47.757/node ($5.97/GPU) | Varies by zone | us-east-1; Capacity Blocks prepaid, price updates next January 2027 | 2026-10-10 | AWS Capacity Blocks |
| CoreWeave | H100 | HGX node, 8 GPU, 128 vCPU, 2 TB RAM | $49.24/node ($6.16/GPU) | Up to 60% off, quote required | $19.71/node ($2.46/GPU) | North America; egress free | 2026-10-10 | CoreWeave pricing |
| Google Cloud | H100 | a3-highgpu-8g, 8 GPU | $88.49/node ($11.06/GPU) | 1-yr $61.38; 3-yr $38.86 | $50.32/node | us-central1; spot price variable | 2026-10-10 | Google accelerator pricing |
| Google Cloud | H100 | a3-megagpu-8g, 8 GPU | $93.40/node ($11.68/GPU) | 1-yr $64.21; 3-yr $40.65 | $53.10/node | us-central1 | 2026-10-10 | Google accelerator pricing |
| Azure | H100 | ND96isr H100 v5, 8 GPU | $98.32/node ($12.29/GPU), via third-party mirror | Not verified | Not verified | East US; official page renders client-side | 2026-10-10 | Azure VM pricing |
| RunPod | H200 | 1 GPU pod, 141 GB, 24 vCPU | $5.29/GPU-hr | Quote required | Not verified | Instant cluster rate $4.31/GPU-hr | 2026-10-10 | RunPod pricing |
| CoreWeave | H200 | HGX node, 8 GPU | $50.44/node ($6.31/GPU) | Quote required | $20.93/node | North America | 2026-10-10 | CoreWeave pricing |
| AWS | H200 | p5e.48xlarge, 8 GPU | Not verified | Capacity Block $54.924/node ($6.87/GPU) | Not verified | US East (Ohio), US West | 2026-10-10 | AWS Capacity Blocks |
| Google Cloud | H200 | a3-ultragpu-8g, 8 GPU | $84.81/node ($10.60/GPU) | 1-yr $58.47; 3-yr $37.21 | $50.83/node | us-central1 | 2026-10-10 | Google accelerator pricing |
| Lambda | B200 SXM6 | 8x instance, 180 GB/GPU | $6.69/GPU-hr ($53.52/node) | 1-Click Cluster, 2 wk–1 yr: $8.87–$9.86/GPU-hr | Not offered | Self-serve | 2026-10-10 | Lambda pricing |
| RunPod | B200 | 1 GPU pod, 180 GB, 28 vCPU | $7.99/GPU-hr | Quote required | Not verified | Per-second billing | 2026-10-10 | RunPod pricing |
| CoreWeave | B200 | HGX node, 8 GPU | $68.80/node ($8.60/GPU) | Quote required | $34.11/node ($4.26/GPU) | North America | 2026-10-10 | CoreWeave pricing |
| Google Cloud | B200 | a4-highgpu-8g, 8 GPU | Not offered on demand | 1-yr $88.93; 3-yr $56.71; DWS Flex-start $64.44 | $39.63/node | us-central1 | 2026-10-10 | Google accelerator pricing |
| AWS | B200 | p6-b200.48xlarge, 8 GPU | Not verified | Capacity Block $113.666/node ($14.21/GPU) | Not verified | us-east-1, us-east-2, us-west-2 | 2026-10-10 | AWS Capacity Blocks |
| Google Cloud | Ironwood (TPU7x) | Per chip | $12.00/chip-hr | 1-yr $8.40; 3-yr $5.40 | Variable | us-central1; London $13.20 | 2026-10-10 | Cloud TPU pricing |
Not verified as publicly available: H200 and B200 on Lambda’s self-serve instance list, B200 on Azure, and any TPU outside Google Cloud. Azure H200 and B200 pricing could not be confirmed from an official page in this pass.
Two things stand out. Google prices the same H100 at nearly four times RunPod’s PCIe rate, but the Google VM bundles 208 vCPUs, 1.87 TB of RAM, 6 TB of local SSD and multi-node networking. And Google sells B200 only through commitments, spot or its Dynamic Workload Scheduler, so there is no plain on-demand rate to compare.
What the hourly GPU cloud price does not tell you
An hourly rate prices access to hardware. It says nothing about how much of that hardware your model will use, or what else the bill will carry.
- Memory decides how many GPUs you rent. A 70-billion-parameter model in FP8 needs roughly 70 GB for weights alone, before the KV cache that holds each conversation’s context. That fits on one 141 GB H200 with room for batching, but is tight on an 80 GB H100, which pushes many teams to two or more cards. The H200’s higher hourly rate can therefore buy a smaller bill.
- Bandwidth sets decode speed. Generating each output token means streaming the model’s weights through the GPU’s memory. AWS describes the B200 instances as offering 1.6x the GPU memory bandwidth of P5en, which matters more for token output than headline FLOPS.
- The node is not just GPUs. CoreWeave’s H100 node includes 128 vCPUs, 2 TB of RAM and 61 TB of local NVMe; a RunPod H100 SXM pod includes 20 vCPUs and 125 GB. Data loading, tokenisation and preprocessing all run on that host side.
- Egress and storage vary sharply. CoreWeave and Lambda advertise no egress fees, while AWS bills data transfer out to the internet beyond a free 100 GB a month. Storage left attached to stopped instances keeps billing: RunPod charges $0.20 per GB-month for idle volume disks, double the running rate.
- Idle time is the largest hidden cost. Google bills a one-minute minimum, then per second, and charges a VM that is idle but RUNNING at full rate. A cluster that waits overnight for the next job costs exactly as much as one serving traffic. As UniverseBlend’s 2026 AI accelerator comparison put it, a rack at 30% use costs about three times as much per token as the same rack at 90%.
On-demand, reserved and spot GPU cloud pricing
- On-demand suits work measured in hours or days: experiments, evaluation runs, short fine-tunes. You pay the highest rate for the right to stop.
- Reservations and commitments pay off when the GPUs would otherwise run continuously. The published discounts are not alike, though. Google’s 1-year commitment on an a3-highgpu-8g cuts the hourly rate from $88.49 to $61.38, and the 3-year rate to $38.86; Google also notes that reserved resources bill whether or not they are running. AWS Capacity Blocks work differently again: you prepay a fixed window up front, at a rate AWS revises on a published schedule. CoreWeave advertises “up to 60%” off for committed use but publishes no rate, so that discount is a ceiling, not a price.
- Spot and preemptible capacity is cheapest on paper and riskiest in practice. CoreWeave lists its spot H100 node at $19.71 an hour, against $49.24 on demand. Google says spot prices can change up to once a day and VMs can be preempted at any time.
Interruption has a measurable cost. Suppose a training job checkpoints every hour and a preemption loses, on average, 30 minutes of work plus 15 minutes to restart. If that happens once every 10 hours, the job does 10 hours of useful work in about 10.75 billed hours. At CoreWeave’s spot rate, that is $19.71 × 10.75 / 10 = $21.19 per useful node-hour, still 57% below on demand. Spot stops paying when preemptions become frequent enough that restart time dominates, or when a deadline cannot absorb the delay. These interruption figures are illustrative assumptions, not measured provider rates.
One caution when comparing discounts across clouds: Google’s committed price, AWS’s prepaid block and CoreWeave’s quoted reservation buy different things, with different cancellation terms and different guarantees of capacity in a specific zone. Compare the effective hourly rate over the full term you would actually commit to.
From GPU cloud pricing to cost per million tokens
The number that matters for inference is what a million output tokens cost. Converting an hourly rate needs only one more figure, sustained output throughput:
\text{Cost per million output tokens} = \frac{\text{hourly compute cost} \times 1{,}000{,}000}{\text{output tokens per second} \times 3{,}600}
As a mathematical illustration only: an instance costing $2.00 an hour that produces 100 output tokens per second costs $2.00 × 1,000,000 / 360,000 = $5.56 per million output tokens.

The throughput evidence
No provider publishes token throughput for its own instances, so this analysis uses MLPerf Inference, the industry benchmark run under audited rules by MLCommons. The workload is Llama 2 70B in the Closed division, Server scenario, which imposes latency limits closer to live serving than the batch-oriented Offline scenario. Throughput is aggregate output tokens per second across all concurrent requests on an eight-GPU system, not the speed any single user sees.
- B200: 102,398 tokens/s on an 8x B200 SXM system (CoreWeave), MLPerf Inference v6.1, reported by NVIDIA, page dated October 2026. Blackwell submissions on this benchmark use FP4 weights.
- H200: 29,228 tokens/s on 8x H200 without Triton, MLPerf Inference v4.1, August 2024, TensorRT-LLM with FP8.
- H100: about 21,500 tokens/s for 8x H100. This is derived, not measured here: NVIDIA’s v4.1 post gives 10,756 tokens/s for one B200 and states that equals 4x the H100’s per-GPU result. 10,756 / 4 × 8 ≈ 21,512.
These figures do not compare like with like. The B200 result is two years newer, uses FP4 rather than FP8, and benefits from two years of software tuning that the Hopper numbers do not capture. Hopper’s cost per token today is probably somewhat lower than shown. The gap is still too wide for that to reverse the ranking.
Results
Table: scroll sideways to view all columns when needed.
| Accelerator / provider | Hourly cost (8-GPU node) | Output throughput | Compute cost per 1M output tokens | Assumptions and source |
|---|---|---|---|---|
| B200 / CoreWeave spot | $34.11 | 102,398 tok/s | $0.09 | MLPerf v6.1 Server; spot, interruptible |
| B200 / Google A4 spot | $39.63 | 102,398 tok/s | $0.11 | As above; Google spot |
| B200 / Lambda | $53.52 | 102,398 tok/s | $0.15 | $6.69 × 8; on demand |
| B200 / CoreWeave | $68.80 | 102,398 tok/s | $0.19 | On demand |
| B200 / AWS Capacity Block | $113.67 | 102,398 tok/s | $0.31 | Prepaid block |
| H200 / CoreWeave spot | $20.93 | 29,228 tok/s | $0.20 | MLPerf v4.1 Server, FP8; spot |
| H200 / RunPod cluster | $34.48 | 29,228 tok/s | $0.33 | $4.31 × 8 |
| H200 / CoreWeave | $50.44 | 29,228 tok/s | $0.48 | On demand |
| H200 / AWS Capacity Block | $54.92 | 29,228 tok/s | $0.52 | p5e; prepaid |
| H200 / Google A3 Ultra | $84.81 | 29,228 tok/s | $0.81 | On demand |
| H100 / Lambda | $31.92 | ~21,512 tok/s (derived) | $0.41 | $3.99 × 8 |
| H100 / AWS Capacity Block | $47.76 | ~21,512 tok/s (derived) | $0.62 | Prepaid block |
| H100 / CoreWeave | $49.24 | ~21,512 tok/s (derived) | $0.64 | On demand |
The B200 node at CoreWeave costs 40% more per hour than the H100 node, yet produces each token for less than a third of the price. That is the whole argument for pricing in tokens: a 4.8x throughput gap overwhelms a 1.4x price gap. It also explains why serving workloads are migrating to newer silicon faster than training, a split UniverseBlend examined in inference chips versus training chips.
These are compute-only estimates. They assume the GPUs run at MLPerf throughput every second of every hour, with prompts and outputs shaped like the benchmark’s OpenOrca dataset (sequences capped at 1,024 tokens). Your prompts, context lengths and latency targets will move the result, and none of these figures include storage, networking, redundancy or engineering time.
Against a hosted API
Together AI lists Llama 3.3 70B at $1.04 per million tokens, input and output alike. That is a newer model than the Llama 2 70B benchmark, so treat the comparison as rough. Even so, the API costs roughly two to eleven times the compute-only estimates above, and buys something real for the difference: no idle hours, no minimum commitment, and someone else handling scaling and failures. Self-hosting wins only if you can keep the hardware busy. A node at 20% utilisation multiplies every figure in the table by five.
A reusable GPU cost calculator
Five inputs are enough to estimate your own cost per million output tokens: hourly instance price, number of instances, measured output tokens per second, productive utilisation, and non-GPU hourly overhead.
- Effective hourly cost = instance hourly cost × number of instances + additional hourly overhead
- Effective token throughput = measured output tokens per second × 3,600 × utilisation fraction
- Cost per million output tokens = effective hourly cost × 1,000,000 / effective token throughput
Utilisation here means the share of paid time spent generating tokens at the measured rate. If your throughput figure already came from a production trace that includes idle gaps, set utilisation to 100% so you do not count the idle time twice.
Worked example. One CoreWeave H200 node at $50.44 an hour, an assumed $2.00 an hour for storage, monitoring and a load balancer, MLPerf throughput of 29,228 tokens per second, and 60% productive utilisation:
- Effective hourly cost: $50.44 × 1 + $2.00 = $52.44
- Effective throughput: 29,228 × 3,600 × 0.60 = 63,132,480 tokens per hour
- Cost: $52.44 × 1,000,000 / 63,132,480 = $0.83 per million output tokens
That is 73% above the $0.48 compute-only figure, from utilisation and a small overhead alone. Real production budgets also carry engineering time, redundancy, retries and orchestration. Treat this as a floor, not a forecast. For a fuller total-cost model that weighs hosted and owned capacity, UniverseBlend’s analysis of AMD’s hybrid AI TCO calculator works through the assumptions behind those vendor tools.
Which provider fits which workload
The published prices support conditional advice rather than a single winner.
- Short experiments and development: single-GPU, per-second pods such as RunPod’s are the cheapest way to get an H100 for an afternoon. Check the host vCPU and RAM before loading a large dataset.
- Fine-tuning and training: multi-GPU jobs need fast GPU-to-GPU links. Full SXM nodes from Lambda, CoreWeave, AWS or Google are the safer choice; a set of separate single-GPU pods is not a substitute.
- Sustained inference: price in tokens. On the evidence above, B200 nodes produced the lowest compute cost per token, and Lambda’s on-demand B200 was the cheapest non-interruptible option checked.
- High-throughput batch generation: spot capacity fits, provided jobs checkpoint. CoreWeave’s and Google’s spot B200 rates gave the lowest token costs in this analysis.
- Bursty workloads: a hosted API or a serverless endpoint usually beats a reserved node, because you stop paying for the troughs.
- Production that cannot tolerate interruption: on-demand or committed capacity only. Google B200 requires a commitment or scheduler-based capacity; AWS publishes prepaid Capacity Blocks.
- Teams needing a specific generation: Google is the only provider here offering TPU7x. AWS lists B200 through Capacity Blocks in nine regions. Lambda’s self-serve list shows B200 and H100 but not H200.
The evidence cannot settle reliability, real capacity in a given week, or support quality. Those need a trial and a written commitment.
How to check a GPU cloud quote before you sign
- Exact instance type and accelerator count, including SXM versus PCIe
- Region and zone, with written confirmation that capacity exists there
- Billing unit and minimum: per second, per minute, per hour or prepaid block
- Included vCPUs, host RAM, local NVMe and attached storage
- Network egress and inter-region transfer charges
- Commitment length, cancellation terms and whether unused hours are billed
- Preemption policy and notice period for spot capacity
- Your own measured throughput on your model, not a vendor benchmark
- Realistic utilisation, including nights and weekends
- Monitoring, orchestration and support costs
Conclusion
GPU cloud pricing is published more openly than at any point in the past three years. Most providers here publish usable rates on static pages; AWS and Azure still keep on-demand rates inside calculators, and Google now lists Ironwood. The problem has shifted from finding a price to interpreting one.
The numbers point one way. The cheapest hour checked was a $2.89 H100 PCIe pod; the cheapest tokens came from an eight-GPU B200 node costing more than $50 an hour. Buyers who compare rental rates will choose the wrong hardware. Measure throughput on your own model, assume honest utilisation, and compare quotes in dollars per million output tokens.
Maintenance note: review this table monthly, and whenever a provider launches a new accelerator, changes a price, or revises billing terms. Record the real verification date for each changed row.
Frequently Asked Questions
How much does it cost to rent an H100 GPU per hour in 2026?
As of October 10, 2026, on-demand H100 prices ran from $2.89 per GPU-hour for a single PCIe pod on RunPod to $11.68 per GPU-hour on a Google Cloud a3-megagpu-8g VM. Full eight-GPU SXM nodes from Lambda and CoreWeave list at $3.99 and $6.16 per GPU-hour.
Is an H200 more expensive than an H100?
Usually, but not everywhere. RunPod charges 33% more ($5.29 versus $3.99 per GPU-hour) and CoreWeave about 2% more per node, while Google’s H200 a3-ultragpu-8g VM ($84.81) is cheaper than its H100 a3-highgpu-8g ($88.49). With 141 GB of memory and higher measured throughput, H200 often costs less per token.
How much does an NVIDIA B200 cost per hour?
Public rates ranged from $6.69 per GPU-hour on Lambda to $14.21 per GPU-hour on an AWS Capacity Block. CoreWeave lists $8.60 on demand and $4.26 on spot. Google offers B200 only through commitments, spot or its Dynamic Workload Scheduler.
Does Google Cloud publish TPU v7 pricing?
Yes. Google’s Cloud TPU pricing page lists Ironwood (TPU7x) at $12.00 per chip-hour on demand in us-central1, $8.40 on a 1-year commitment and $5.40 on a 3-year commitment. Prices are per chip, not per VM, and are not directly comparable with GPU-hours.
Is spot GPU pricing worth the interruption risk?
For checkpointed training and batch jobs, usually yes: CoreWeave’s spot H100 node costs 60% less than on demand, leaving room to absorb lost work. For latency-sensitive production serving, no, because a preemption becomes an outage.
Is renting a GPU cheaper than using an AI API?
Only at high utilisation. Compute-only costs on rented 70B-class serving ranged from about $0.09 to $0.81 per million output tokens at full benchmark load, against $1.04 for Llama 3.3 70B on Together AI’s API. Below roughly 20% to 60% utilisation, depending on hardware and rate, the API is often cheaper.
How do I calculate GPU inference cost per million tokens?
Multiply the hourly cost by 1,000,000, then divide by output tokens per second × 3,600 × your utilisation fraction. Use throughput measured on your own model and prompt lengths.
What additional costs should I include in a GPU cloud budget?
Storage, data egress, idle time, host CPU and memory, monitoring, orchestration, redundancy and engineering time. Idle time is usually the largest, because most providers bill a running instance whether or not it does useful work.
Keep reading
Here are the latest posts from the blog.

GPU Cloud Pricing 2026: What an Hour Actually Costs

AI API Rate Limits in 2026: The Numbers That Bind

The $60B AI Chip Financing and Its Residual-Value Problem
