On-demand
Pay for capacity when it is running.
Good fit: Use it for the first benchmark, debugging, or work that cannot be interrupted.
Check: Watch hourly rate, startup time, region, and attached storage.
Cloud GPU rental guide
Cloud GPU rental lets you access accelerator capacity without buying a server. The practical choice is not simply the lowest GPU-hour: match memory to the workload, normalize the complete instance, run a small benchmark, and choose the rental model that keeps useful work reliable.
Provider rows on this page were last refreshed Aug 8, 2026. Availability, rates, and billing terms change, so verify the final offer before a long run.
Short answer
Cloud GPU rental is temporary access to a GPU-backed virtual machine, bare-metal host, marketplace listing, or managed endpoint. You pay for the capacity you use instead of owning the accelerator, rack, power system, cooling, and replacement parts. That makes rental useful when demand is uncertain, a project is short, or you need to test several GPU generations before choosing a long-term platform.
Start with the cloud GPU pricing comparison for current provider rows, then use this page to decide which billing model and provider type fits. For a step-by-step AI setup, see the how to rent a GPU for AI guide; it covers launch workflow rather than the broader rental-service decision.
The rental model changes the economics
“Cloud GPU rental” covers several operating models. A low marketplace spot listing and a reserved multi-GPU node are not substitutes even when both advertise an hourly GPU rate. Decide how much interruption, setup work, and capacity risk the workload can absorb before comparing providers.
On-demand
Good fit: Use it for the first benchmark, debugging, or work that cannot be interrupted.
Check: Watch hourly rate, startup time, region, and attached storage.
Spot or interruptible
Good fit: Use it for checkpointed training, batch jobs, rendering, and experiments that can resume.
Check: Budget for preemption, retries, queue time, and lost progress.
Dedicated or committed
Good fit: Use it for steady production inference, team capacity, data locality, or multi-GPU work.
Check: Check contract length, setup fees, support, network, and exit terms.
Serverless or autoscaled
Good fit: Use it for bursty inference when an always-on instance would sit idle.
Check: Compare cold starts, warm workers, platform fees, concurrency, and latency.
Use live data as a first filter
This table uses GPU Cost's current D1 pricing rows. It shows the lowest tracked on-demand rate and a simple 730-hour always-on checkpoint, which is useful for screening but not a quote. It excludes storage, CPU, RAM, network transfer, taxes, platform fees, and interrupted or idle work.
| GPU | Lowest tracked rate | 730-hour checkpoint | Providers |
|---|---|---|---|
| RTX 309024 GB VRAM | $0.015/GPU-hour | $10.95/month | 3 |
| V10016 GB VRAM | $0.025/GPU-hour | $18.54/month | 3 |
| L422 GB VRAM | $0.032/GPU-hour | $23.65/month | 4 |
| RTX A400016 GB VRAM | $0.060/GPU-hour | $43.80/month | 4 |
| RTX 408015 GB VRAM | $0.098/GPU-hour | $71.69/month | 2 |
| RTX 30708 GB VRAM | $0.130/GPU-hour | $94.90/month | 1 |
| Tesla V10032 GB VRAM | $0.140/GPU-hour | $102.20/month | 5 |
| RTX 308010 GB VRAM | $0.170/GPU-hour | $124.10/month | 1 |
| RTX 3080 Ti12 GB VRAM | $0.180/GPU-hour | $131.40/month | 1 |
| V100 FHHL16 GB VRAM | $0.190/GPU-hour | $138.70/month | 1 |
| Tesla V10016 GB VRAM | $0.190/GPU-hour | $138.70/month | 1 |
| V100 SXM216 GB VRAM | $0.230/GPU-hour | $167.90/month | 1 |
Data refresh: Aug 8, 2026. For provider filters and all tracked rows, open cloud GPU pricing. For utilization and overhead assumptions, use the GPU rental cost calculator.
Normalize the offer before ranking providers
Provider type predicts the tradeoff, but the individual offer still matters. A hyperscaler may fit an existing network and identity stack; a specialized GPU cloud may provision faster; a marketplace may expose a lower starting price with more host variation. Compare the same GPU, memory, region, storage, and interruption policy across at least two options.
| Provider | Type | Lowest on-demand | Lowest spot | GPU coverage |
|---|---|---|---|---|
| Vast.ai | GPU marketplace | Check offers | — | — |
| Genesis Cloud | Specialized GPU cloud | Check offers | — | — |
| Massed Compute | GPU marketplace | Check offers | — | — |
| Vast.ai | GPU marketplace | $0.015/GPU-hour | $0.014 | 8 |
| TensorDock | GPU marketplace | $0.060/GPU-hour | — | 9 |
| RunPod | Specialized GPU cloud | $0.130/GPU-hour | $0.100 | 50 |
| Datacrunch (Verda) | Specialized GPU cloud | $0.140/GPU-hour | — | 9 |
| Genesis Cloud | Specialized GPU cloud | $0.250/GPU-hour | — | 4 |
| Google Cloud Platform | Hyperscaler | $0.350/GPU-hour | $0.140 | 10 |
| CoreWeave | Specialized GPU cloud | $0.390/GPU-hour | — | 12 |
| Amazon Web Services | Hyperscaler | $0.420/GPU-hour | $0.182 | 8 |
| Microsoft Azure | Hyperscaler | $0.454/GPU-hour | — | 6 |
| Jarvis Labs | Specialized GPU cloud | $0.490/GPU-hour | — | 6 |
| Lambda Labs | Specialized GPU cloud | $0.690/GPU-hour | — | 13 |
| Paperspace | Specialized GPU cloud | $0.760/GPU-hour | — | 7 |
| Oracle Cloud | Hyperscaler | $1.00/GPU-hour | — | 3 |
| Fluidstack | Specialized GPU cloud | $1.30/GPU-hour | — | 5 |
A selection workflow that survives a real bill
The first offer that launches is not always the right long-term choice. Use a small, reversible test to separate a low advertised rate from a low cost per useful result. This workflow is deliberately different from a provider directory: it focuses on the evidence you need before scaling.

Record model size, precision, context length, batch size, input data, expected concurrency, and the result you need to measure.
Choose the smallest GPU that fits weights, KV cache, optimizer state, activations, and headroom. A fast GPU that cannot fit is not a usable bargain.
Match GPU variant, CPU, RAM, local disk, persistent storage, region, network, billing unit, and interruption policy before comparing rates.
Start on-demand, validate drivers and data paths, then measure completed tokens, images, samples, or training steps per dollar.
Stop idle capacity, remove unused volumes, set alerts, checkpoint spot jobs, and record the effective cost of retries and queue time.
The number a price table cannot answer
Compare cost per useful result, not only cost per hour. A faster GPU can be cheaper when it finishes a run sooner; a low marketplace rate can be expensive when a job is interrupted, a volume remains attached, or data has to cross regions. For request traffic, include warm capacity and platform fees instead of treating every second as active GPU work.
GPU compute + storage + CPU/RAM + network + idle time + retries + platform/support feesGPU rate, number of GPUs, active hours, and whether billing is per GPU, per instance, per second, or per request.
Persistent volumes, snapshots, object storage, and checkpoint copies can continue billing after the GPU is stopped.
CPU, RAM, local NVMe, networking, and the complete node shape can change the cost more than the headline accelerator rate.
Dataset, model-weight, and output transfers may add regional egress or cross-service charges.
Spot preemption, queue time, restart work, failed runs, and support time are costs when the job is not fully resumable.
Images, drivers, secrets, monitoring, backups, access controls, and incident response affect the useful cost of a service.
For a personalized estimate, use the GPU rental cost calculator. If the workload is intermittent, compare a dedicated instance with the serverless GPU pricing calculator.
Use the GPU server cost guide and GPU server cost calculator for the ownership side. The correct break-even point compares equivalent completed work over the same period, not a monthly GPU-hour total by itself.
Cloud GPU rental gives you temporary access to a GPU instance through a cloud provider instead of purchasing and operating the hardware. You normally pay for active compute time plus storage, CPU or RAM, network transfer, and any platform fees.
Normalize the comparison around the same GPU memory, CPU and RAM, region, storage, billing unit, and interruption policy. Then estimate the cost of useful completed work rather than comparing a headline GPU-hour alone.
Rental is usually easier to justify for experiments, bursty demand, changing GPU requirements, and teams without a rack or operations staff. Buying can win at sustained utilization when power, cooling, support, downtime, and refresh risk are included in the ownership model.
Marketplace or spot capacity may show the lowest rate, but the cheapest useful option is the one that finishes the workload reliably after storage, queue time, interruptions, retries, and data transfer are counted.
Use on-demand for a first benchmark or work that cannot be interrupted. Use spot when the job checkpoints and can resume. Choose dedicated or committed capacity when availability, predictable performance, data locality, or long-running utilization matters more than the lowest list price.
Multiply the hourly rate by the hours the GPU is actually running, then add storage, CPU or RAM, network, idle time, retries, and support. A 730-hour calculation is an always-on checkpoint, not a recommendation that every workload should run continuously.
Yes. Rental works well for development, batch inference, evaluation, and variable traffic. For an API with intermittent requests, compare an always-on instance with serverless or autoscaled GPU capacity so warm workers and idle time are visible in the estimate.
Use these first-party destinations to verify the final offer after narrowing the GPU, region, billing model, and workload:
These links are verification references, not a guarantee that one provider is always cheapest. GPU Cost's live rows are a comparison starting point; confirm region, instance shape, storage, network, and billing terms before a long run.