Cloud GPU rental guide

Cloud GPU Rental: Pricing, Providers, and How to Choose

Cloud GPU rental lets you access accelerator capacity without buying a server. The practical choice is not simply the lowest GPU-hour: match memory to the workload, normalize the complete instance, run a small benchmark, and choose the rental model that keeps useful work reliable.

Provider rows on this page were last refreshed Aug 8, 2026. Availability, rates, and billing terms change, so verify the final offer before a long run.

Editorial illustration of GPU servers connected to cloud computing capacity and a cost meter
Cloud GPU rental turns a hardware purchase into a workload-sized commitment.

Short answer

What is cloud GPU rental?

Cloud GPU rental is temporary access to a GPU-backed virtual machine, bare-metal host, marketplace listing, or managed endpoint. You pay for the capacity you use instead of owning the accelerator, rack, power system, cooling, and replacement parts. That makes rental useful when demand is uncertain, a project is short, or you need to test several GPU generations before choosing a long-term platform.

Start with the cloud GPU pricing comparison for current provider rows, then use this page to decide which billing model and provider type fits. For a step-by-step AI setup, see the how to rent a GPU for AI guide; it covers launch workflow rather than the broader rental-service decision.

The rental model changes the economics

Choose on-demand, spot, dedicated, or serverless GPU capacity

“Cloud GPU rental” covers several operating models. A low marketplace spot listing and a reserved multi-GPU node are not substitutes even when both advertise an hourly GPU rate. Decide how much interruption, setup work, and capacity risk the workload can absorb before comparing providers.

On-demand

Pay for capacity when it is running.

Good fit: Use it for the first benchmark, debugging, or work that cannot be interrupted.

Check: Watch hourly rate, startup time, region, and attached storage.

Spot or interruptible

Trade availability for a lower price.

Good fit: Use it for checkpointed training, batch jobs, rendering, and experiments that can resume.

Check: Budget for preemption, retries, queue time, and lost progress.

Dedicated or committed

Reserve a host, node, or term for predictable access.

Good fit: Use it for steady production inference, team capacity, data locality, or multi-GPU work.

Check: Check contract length, setup fees, support, network, and exit terms.

Serverless or autoscaled

Pay around requests, runtime, or warm capacity.

Good fit: Use it for bursty inference when an always-on instance would sit idle.

Check: Compare cold starts, warm workers, platform fees, concurrency, and latency.

Use live data as a first filter

Cloud GPU rental pricing checkpoints

This table uses GPU Cost's current D1 pricing rows. It shows the lowest tracked on-demand rate and a simple 730-hour always-on checkpoint, which is useful for screening but not a quote. It excludes storage, CPU, RAM, network transfer, taxes, platform fees, and interrupted or idle work.

GPULowest tracked rate730-hour checkpointProviders
RTX 309024 GB VRAM$0.015/GPU-hour$10.95/month3
V10016 GB VRAM$0.025/GPU-hour$18.54/month3
L422 GB VRAM$0.032/GPU-hour$23.65/month4
RTX A400016 GB VRAM$0.060/GPU-hour$43.80/month4
RTX 408015 GB VRAM$0.098/GPU-hour$71.69/month2
RTX 30708 GB VRAM$0.130/GPU-hour$94.90/month1
Tesla V10032 GB VRAM$0.140/GPU-hour$102.20/month5
RTX 308010 GB VRAM$0.170/GPU-hour$124.10/month1
RTX 3080 Ti12 GB VRAM$0.180/GPU-hour$131.40/month1
V100 FHHL16 GB VRAM$0.190/GPU-hour$138.70/month1
Tesla V10016 GB VRAM$0.190/GPU-hour$138.70/month1
V100 SXM216 GB VRAM$0.230/GPU-hour$167.90/month1

Data refresh: Aug 8, 2026. For provider filters and all tracked rows, open cloud GPU pricing. For utilization and overhead assumptions, use the GPU rental cost calculator.

Normalize the offer before ranking providers

How to compare cloud GPU rental providers

Provider type predicts the tradeoff, but the individual offer still matters. A hyperscaler may fit an existing network and identity stack; a specialized GPU cloud may provision faster; a marketplace may expose a lower starting price with more host variation. Compare the same GPU, memory, region, storage, and interruption policy across at least two options.

ProviderTypeLowest on-demandLowest spotGPU coverage
Vast.aiGPU marketplaceCheck offers
Genesis CloudSpecialized GPU cloudCheck offers
Massed ComputeGPU marketplaceCheck offers
Vast.aiGPU marketplace$0.015/GPU-hour$0.0148
TensorDockGPU marketplace$0.060/GPU-hour9
RunPodSpecialized GPU cloud$0.130/GPU-hour$0.10050
Datacrunch (Verda)Specialized GPU cloud$0.140/GPU-hour9
Genesis CloudSpecialized GPU cloud$0.250/GPU-hour4
Google Cloud PlatformHyperscaler$0.350/GPU-hour$0.14010
CoreWeaveSpecialized GPU cloud$0.390/GPU-hour12
Amazon Web ServicesHyperscaler$0.420/GPU-hour$0.1828
Microsoft AzureHyperscaler$0.454/GPU-hour6
Jarvis LabsSpecialized GPU cloud$0.490/GPU-hour6
Lambda LabsSpecialized GPU cloud$0.690/GPU-hour13
PaperspaceSpecialized GPU cloud$0.760/GPU-hour7
Oracle CloudHyperscaler$1.00/GPU-hour3
FluidstackSpecialized GPU cloud$1.30/GPU-hour5
Shortlist rule: keep one offer from a hyperscaler or enterprise cloud, one from a specialized GPU cloud, and one marketplace or spot option when the workload can tolerate interruption. The comparison is more useful when every row can run the same benchmark.

A selection workflow that survives a real bill

How to choose a cloud GPU rental

The first offer that launches is not always the right long-term choice. Use a small, reversible test to separate a low advertised rate from a low cost per useful result. This workflow is deliberately different from a provider directory: it focuses on the evidence you need before scaling.

Editorial flow showing a GPU, cloud server, launch step, storage volume, and cost monitor
Choose the workload fit first, then validate the provider and the complete cost.
1

Write the workload spec

Record model size, precision, context length, batch size, input data, expected concurrency, and the result you need to measure.

2

Set the memory floor

Choose the smallest GPU that fits weights, KV cache, optimizer state, activations, and headroom. A fast GPU that cannot fit is not a usable bargain.

3

Normalize providers

Match GPU variant, CPU, RAM, local disk, persistent storage, region, network, billing unit, and interruption policy before comparing rates.

4

Run a reversible benchmark

Start on-demand, validate drivers and data paths, then measure completed tokens, images, samples, or training steps per dollar.

5

Control the bill

Stop idle capacity, remove unused volumes, set alerts, checkpoint spot jobs, and record the effective cost of retries and queue time.

The number a price table cannot answer

Cloud GPU rental cost is more than the GPU-hour

Compare cost per useful result, not only cost per hour. A faster GPU can be cheaper when it finishes a run sooner; a low marketplace rate can be expensive when a job is interrupted, a volume remains attached, or data has to cross regions. For request traffic, include warm capacity and platform fees instead of treating every second as active GPU work.

Useful workload costGPU compute + storage + CPU/RAM + network + idle time + retries + platform/support fees
01

Compute

GPU rate, number of GPUs, active hours, and whether billing is per GPU, per instance, per second, or per request.

02

Storage

Persistent volumes, snapshots, object storage, and checkpoint copies can continue billing after the GPU is stopped.

03

Host shape

CPU, RAM, local NVMe, networking, and the complete node shape can change the cost more than the headline accelerator rate.

04

Network

Dataset, model-weight, and output transfers may add regional egress or cross-service charges.

05

Reliability

Spot preemption, queue time, restart work, failed runs, and support time are costs when the job is not fully resumable.

06

Operations

Images, drivers, secrets, monitoring, backups, access controls, and incident response affect the useful cost of a service.

For a personalized estimate, use the GPU rental cost calculator. If the workload is intermittent, compare a dedicated instance with the serverless GPU pricing calculator.

When cloud GPU rental beats buying a server

Rent first when

  • Demand is uncertain, seasonal, or bursty.
  • You are testing model size, GPU generation, framework, or deployment shape.
  • You need capacity quickly without procurement, power, cooling, and hardware support.
  • You need to compare several GPU types before choosing a long-term platform.

Buying may win when

  • Utilization is high, predictable, and likely to remain useful for years.
  • Data locality, compliance, or latency requires local control.
  • You already operate the rack, network, power, cooling, and monitoring stack.
  • The ownership model includes realistic depreciation, support, downtime, and refresh risk.

Use the GPU server cost guide and GPU server cost calculator for the ownership side. The correct break-even point compares equivalent completed work over the same period, not a monthly GPU-hour total by itself.

Cloud GPU rental FAQ

What is cloud GPU rental?

Cloud GPU rental gives you temporary access to a GPU instance through a cloud provider instead of purchasing and operating the hardware. You normally pay for active compute time plus storage, CPU or RAM, network transfer, and any platform fees.

How do I compare cloud GPU rental prices?

Normalize the comparison around the same GPU memory, CPU and RAM, region, storage, billing unit, and interruption policy. Then estimate the cost of useful completed work rather than comparing a headline GPU-hour alone.

Is cloud GPU rental cheaper than buying a GPU server?

Rental is usually easier to justify for experiments, bursty demand, changing GPU requirements, and teams without a rack or operations staff. Buying can win at sustained utilization when power, cooling, support, downtime, and refresh risk are included in the ownership model.

What is the cheapest cloud GPU rental option?

Marketplace or spot capacity may show the lowest rate, but the cheapest useful option is the one that finishes the workload reliably after storage, queue time, interruptions, retries, and data transfer are counted.

Should I use on-demand, spot, or dedicated GPU rental?

Use on-demand for a first benchmark or work that cannot be interrupted. Use spot when the job checkpoints and can resume. Choose dedicated or committed capacity when availability, predictable performance, data locality, or long-running utilization matters more than the lowest list price.

How much does a cloud GPU rental cost per month?

Multiply the hourly rate by the hours the GPU is actually running, then add storage, CPU or RAM, network, idle time, retries, and support. A 730-hour calculation is an always-on checkpoint, not a recommendation that every workload should run continuously.

Can I use cloud GPU rental for AI inference?

Yes. Rental works well for development, batch inference, evaluation, and variable traffic. For an API with intermittent requests, compare an always-on instance with serverless or autoscaled GPU capacity so warm workers and idle time are visible in the estimate.

Official provider references

Use these first-party destinations to verify the final offer after narrowing the GPU, region, billing model, and workload:

These links are verification references, not a guarantee that one provider is always cheapest. GPU Cost's live rows are a comparison starting point; confirm region, instance shape, storage, network, and billing terms before a long run.