Quick Answer
Cloud GPU pricing in 2026 spans roughly a 5x spread for identical hardware. Including spot and marketplace rates, H100s rent from $1.49/hour to $6.98, A100s from $0.68 to $5.03, L4s from $0.13 to $0.80, and B200s from $3.99 to $16.11. Dedicated GPU clouds average about half of hyperscaler rates.
Here’s the strangest story in cloud pricing right now. AI API prices crashed up to 80% this summer in an open price war. GPU rental prices didn’t move: GetDeploying’s index, built from 76 providers and 98 GPU models, has on-demand rates up just 1.7% over the past 12 months. Both are AI compute. One collapsed in price, one held firm while Blackwell demand absorbed supply. If you rent GPUs, the market just told you something: the discount is not coming from the rate card.
Which makes provider choice the biggest lever you have, and the least examined. The same H100, the same hour of work, costs $1.49 on one platform and $6.98 on another, and in CloudZero’s 2026 AI ROI survey of 260 finance leaders, 34% couldn’t produce a credible ROI number for their AI spend. Hard to prove a return when the same input costs 5x depending on where you bought it and nobody’s tracking which. This page is the map.
How much does a cloud GPU cost by provider?
The pricing is per GPU-hour, on-demand:
| Provider | H100 80GB | A100 80GB | L4 | B200 | Type |
|---|---|---|---|---|---|
| Vast.ai | 1.49â2.27 | from $0.68 (spot) | from $0.13 (spot) | POA | Marketplace |
| RunPod | 1.99â2.99 | TBC | TBC | POA | Dedicated cloud |
| Lambda | 3.29â3.99 | TBC | â | $4.99 | Dedicated cloud |
| CoreWeave | 4.25â6.16 | 2.21â2.70 | â | ~$5.50 | Dedicated cloud |
| GCP | from ~$2.25 (spot) | 5.03 (3.67 for 40GB) | $0.70 | up to $16.11 (A4) | Hyperscaler |
| AWS | ~$6.88 (p5 list) | TBC (p4d) | TBC (G6) | ~$9.36 (capacity blocks) | Hyperscaler |
| Azure | up to $6.98 (ND) | TBC (NC/ND) | TBC (NV) | POA | Hyperscaler |
| Market median | $4.19 clouds / $7.89 hyperscalers | $2.00 / $3.67 | TBC | $7.88 / $15.18 | GetDeploying |
Before the patterns, the finance framing that should sit over this whole table: these are prices, not costs. Your cost is the rate times the hours times the fraction of hours that produced anything, and only the first number is printed here.
Now, three patterns worth more than any single rate:
- First, the hyperscaler premium: GetDeploying’s index puts median on-demand H100 pricing at $4.19/hour on dedicated GPU clouds versus $7.89 on hyperscalers, an 88% premium for identical silicon. The gap widens on newer parts: 93% on B200s and 132% on H200s.
- Second, spot changes everything: GCP spot cuts 60% to 91% off on-demand, and marketplace spot occasionally drops H100s under $2.
- Third, the spread widens as GPUs age: L4 rates run from $0.13 spot to $0.80 on-demand, a 6x range on a $2,000 card.
Report
Finance needs to prove AIâs return: CloudZero report
260 senior finance leaders (more than half CFOs) told us why the speed of seeing AI spend, not the size of it, separates who pulls ahead on AI from who gets burned.
What is GPU as a service?
GPU as a service means exactly what it sounds like: someone else owns the silicon, runs the data center, and rents you compute by the hour, minute, or second. You bring the workload and the credit card.
The GPU cloud providers market splits three ways: hyperscalers (AWS, Azure, GCP), dedicated GPU clouds (CoreWeave, Lambda, RunPod), and marketplaces (Vast.ai) that aggregate third-party hardware.
Whatever you call it (GPU rental, remote GPU access, GPU cloud computing, renting a GPU cloud server), the model won because AI workloads are spiky: training runs need 64 GPUs for a week, then zero for a month, and hourly access fits that shape in a way ownership never did.
The catch is the same as every rental: convenience is priced in, and idle rented hours cost exactly as much as busy ones. Teams that forget that second sentence fund the entire industry’s margins.
How much does CoreWeave cost?
CoreWeave pricing is published and more transparent than it used to be: H100 PCIe at $4.25/hour, an 8x HGX H100 node at $49.24/hour (about $6.16 per GPU), A100 80GB from $2.21/hour single or $2.70 per GPU on 8x systems, and B200 capacity around $5.50/hour. Reserved contracts cut on-demand rates by up to 60%.
CoreWeave took a $2 billion Nvidia investment in January 2026 to accelerate a buildout of more than 5 gigawatts of AI capacity by 2030, and became the reference AI-native cloud the dedicated tier is measured against.
The CoreWeave H100 pricing per hour runs above Lambda and RunPod, and that’s not an accident: you’re paying for InfiniBand networking at 400 Gbps, bare-metal Kubernetes, and the cluster-scale interconnect that distributed training at 100+ GPUs actually requires. For single-GPU work, CoreWeave is expensive. For a 512-GPU training run, the interconnect premium pays for itself in wall-clock time, which is the only clock that bills.
Lambda, RunPod, and Vast: the dedicated cloud tier
Lambda Labs GPU cloud pricing runs $3.29 to $3.99/hour for H100s (SXM at the top) and $4.99 for B200s, with A100s in the dedicated-tier range below CoreWeave’s $2.21 anchor. The differentiator that matters more than any rate: no egress fees. On data-heavy jobs, hyperscaler egress can exceed the GPU line itself, so Lambda’s zero is a real discount hiding outside the rate card.
RunPod is the price aggressor: H100s from $1.99/hour on community tier ($2.89-$2.99 secure), per-second billing, and consumer GPUs (an RTX 4090 from $0.34/hour) for workloads that don’t need data center silicon. Community tier means shared, multi-tenant infrastructure, which is fine for experiments and rough for compliance.
Vast.ai runs the GPU rental marketplace model to its logical floor: verified-host H100s at $1.49 to $2.27/hour, and the cheapest listed GPU on the internet (an RTX 4070 at about $0.016/hour). Quality varies by host because the hosts are the product; treat it as the spot market it is.
Most inference shoppers arrive with the same request: something cheap that can serve a 7B model. For that job, an L4 at $0.44 to $0.80/hour beats an A100 at 3-5x the price, and the skill is matching the GPU to the job, not maxing the GPU.
Google Cloud GPU pricing: the full ladder
Google Cloud GPU pricing is the most legible of the hyperscalers, because GPU rates sit inside standard Compute Engine pricing, attached to machine families. The catalog spans every generation: B200 180GB on A4 instances at up to $16.11 per GPU on-demand (the price of not negotiating), H100 and H200 on A3, A100 at $3.67 (40GB) and $5.03 (80GB) on A2, L4 at $0.70 on G2, and T4 from $0.54 on N1 for legacy work.
GCP GPU pricing has two levers that beat any list rate: spot instances at 60% to 91% off for interruptible work, and the Dynamic Workload Scheduler, which queues batch jobs for reduced-rate capacity without spot’s interruption risk.
The Google Cloud L4 GPU price per hour at $0.70 on-demand makes it the default serving GPU for small and mid-size models. Market-wide, L4s run $0.44 to $0.80, with a median that fell 7% over the past year. Down-ladder, GCP T4 pricing from $0.54 covers legacy inference, and the GCP A100 cost at $3.67/$5.03 sits mid-pack; up-ladder, GCP H100 pricing competes mainly through spot.
The quota catch nobody budgets for: GCP GPU quotas start at zero. Approval takes hours to days, which is only a problem if you discover it the day the training run was scheduled.
And the part no pricing page covers: a discounted rate on an unallocated invoice is still an unexplainable number at the board meeting.
What do AWS and Azure charge for GPUs?
AWS GPU pricing anchors high: the 8x H100 p5.48xlarge lists at $55.04/hour (about $6.88 per GPU), with effective rates near $4 through capacity blocks and commitments (the same discount discipline covered in our AWS cost optimization guide), and B200 capacity blocks around $9.36 per GPU-hour. Azure H100 pricing tops the on-demand market at up to $6.98/hour on the ND-series, a premium that extends to its A100 tiers.
What the hyperscalers sell isn’t the GPU; it’s the GPU inside your existing security perimeter, IAM, and data gravity. That’s worth a premium. Whether it’s worth nearly double, per the market medians, is a question each workload should answer rather than inherit.
What’s the cheapest GPU cloud?
The cheapest option depends on what “cheap” has to survive. If you’re shopping H100s and A100s specifically, the floors are $1.49 and $0.68 per hour:
- Vast.ai for absolute floor prices ($1.49 H100s, cents-per-hour consumer cards), if you can tolerate host variability.
- RunPod for the best cheap-but-managed balance: $1.99 H100s, per-second billing, real orchestration.
- Salad and TensorDock for community-powered consumer GPU pools at deep discounts, suited to stateless, fault-tolerant work.
- GCP spot for the cheapest hyperscaler route: 60-91% off inside an enterprise perimeter.
- Lambda for the lowest total cost on data-heavy jobs, because zero egress beats a lower hourly rate once the dataset is large.
The honest footnote: every list above is one price change from stale, which is why the change log below exists.
And because “cheapest” is workload-dependent, the decision table:
| Your workload | Best-fit provider tier | Why |
|---|---|---|
| Experiments, fine-tuning, bursty jobs | RunPod, Vast.ai | Per-second billing, spot floors, easy teardown |
| Sustained single-node training | Lambda | Competitive H100 rates, zero egress on big datasets |
| Multi-node distributed training (100+ GPUs) | CoreWeave | InfiniBand interconnect; wall-clock savings beat the rate premium |
| Production inference, small models | GCP (L4 on G2) | $0.70/hour serving tier, spot and DWS discounts |
| Regulated or data-gravity workloads | AWS, Azure | The GPU inside your existing perimeter; premium priced in |
| Stateless, fault-tolerant batch | Salad, TensorDock | Community GPU pools at the market floor |
Rent smart: the questions that beat the rate card
Provider choice is lever one. These are levers two through four, and they’re usually bigger:
Spot vs on-demand. If the workload checkpoints and restarts, spot’s 60-91% discount is free money. If it doesn’t, one interruption eats a week of savings.
The costs outside the rate card. Egress is the big one: moving a large dataset out of a hyperscaler can cost more than the GPU hours that processed it, which is why Lambda’s zero-egress policy and marketplace download fees belong in every comparison. Storage attached to idle instances keeps billing after the GPUs stop. And quota lead times (GCP GPU quotas start at zero) are a cost paid in schedule instead of dollars, invoiced the day your training run can’t start. Egress, idle storage, and quota delay are all invoiced separately from the GPU-hour, or not invoiced at all and paid in schedule. Add them up before you compare two providers on rate alone.
The GPU-to-job match. Serving a 7B model on an A100 is paying for VRAM you’ll never touch. The L4-vs-A100 decision is a 3-5x cost decision that most teams never consciously make:
| GPU | Hourly range | Right-sized for |
|---|---|---|
| L4 (24GB) | $0.44-$0.80 | Serving models to ~13B params, video, light training |
| A100 80GB | $0.68 (spot)-$5.03 | Mid-size training, larger model serving |
| H100 80GB | $1.49-$6.98 | Serious training, high-throughput inference |
| B200 (180GB) | $3.99-$9.36+ | Frontier training; ~2.5x H100 throughput often wins per result |
Reading the table bottom-up is the cost trick: the most expensive GPU per hour is frequently the cheapest per result on training, and the cheapest GPU per hour is usually the right answer for serving. Hourly rate is a price; cost per result is the bill.
Sustained high utilization eventually favors owning: a used 8x H100 system pays back against rentals in roughly seven months at full utilization, and closer to two years at 30% (the full math is in our H100 price guide). Plenty of inference workloads shouldn’t touch a GPU at all when a per-token API costs 10x less, priced in the LLM API pricing comparison. The full decision framework, including what AI actually costs end to end, is in the how much does AI cost guide and the inference cost breakdown.
How do you track GPU spend across all these providers?
Now the part every provider’s pricing page skips: the multi-provider bill. Real teams don’t pick one row from the master table; they end up with three. Training on Lambda because of egress, burst experiments on RunPod because of per-second billing, production inference on AWS because that’s where the data lives.
That means three invoices, three billing units (GPU-hours, GPU-seconds, instance-hours), zero shared labels, and a finance leader trying to answer “what does our AI infrastructure cost per model” from PDFs that don’t even agree on what an hour is.
That’s the specific job CloudZero’s GPU story is built around. Native AWS, Azure, and GCP integrations pull hyperscaler GPU spend automatically, and the AnyCost framework ingests dedicated GPU cloud and marketplace bills into the same normalized stream, so the AI Hub shows one view of GPU spend regardless of who invoiced it.
From there, Dimensions allocate that spend to cost per training run, per model, and per team without tagging. Budgets and forecasting turn commodity-volatile GPU rates into forward numbers a board can hold. And anomaly detection with hour-level granularity catches the cluster someone forgot to spin down while it’s a Tuesday problem instead of a month-end one.
And here is one more thing: the market will keep handing you cheaper GPU hours, and none of it matters if you can’t say what the hours produced. Rate optimization is table stakes; GPU cloud ROI is the game, and the spend data across the industry (in CloudZero’s State of AI Costs report and the AI spend management guide) says most teams are still playing the first one.
It works for teams shaped exactly like this page’s reader: Philip Tsai, VP of engineering at Helm.ai, put it this way: “CloudZero not only allows us to analyze and track costs efficiently but also ensures that our innovations remain sustainable as we grow.” Sustainability is the operative word. A lean team on a volatile compute profile doesn’t need a lower hourly rate so much as it needs to know which experiments are earning their keep before the next funding conversation.
Schedule a demo to see your multi-provider GPU spend allocated to cost per model, per training run, and per team, including the idle hours no rate card shows you. More stories from teams like Helm.ai on the customers page.