Data Center NVIDIA

L4

Ada Lovelace Architecture · 22GB GDDR6 · PCIe

VRAM
22GB
FP16
60.6
TDP
72W
Hardware Price
$2.8k
MSRP: $2.5k
Cloud from
$0.032/hr
11 providers
Cheapest at Vast.ai →

Quick Insights

Performance/Dollar
21.64 TFLOPS/$k
FP16 performance per $1000
VRAM/Dollar
7.9 GB/$k
VRAM per $1000
vs Data Center Average
-66% perf
FP16 TFLOPS comparison
Cloud Availability
11 providers
from $0.032/hr

Specifications

VRAM 22GB GDDR6
Memory Bandwidth 300 GB/s
FP16 TFLOPS 60.6
Tensor TFLOPS 242.0
FP32 TFLOPS 30.3
TDP 72W
Form Factor PCIe
Architecture Ada Lovelace
NVLink No
Release Date 2023-03

L4 Price and Rental Notes

NVIDIA L4 price searches include used-card pricing, Google Cloud L4 pricing, 24GB VRAM checks, and L4 vs T4 comparisons. The main opportunity is to clarify when L4 is cheaper than larger GPUs and when it is a practical T4 upgrade.

NVIDIA L4 price for cloud inference

L4 is usually evaluated for efficient inference, video workloads, and smaller AI models. Compare cloud hourly price, VRAM, and availability against A10, L40S, and RTX 4090 options.

Google Cloud L4 pricing context

Google Cloud is a common L4 reference point, but final cost depends on region, machine type, committed use discounts, and attached CPU or memory. Use per-GPU pricing only as the first filter.

L4 cost per inference example

Suppose an L4 instance costs $0.80 per active hour and serves 180,000 requests in that hour. The GPU-and-instance cost is about $4.44 per million requests before storage, networking, orchestration, and idle capacity. Use your measured throughput instead of assuming the lowest hourly quote produces the lowest inference cost.

Best fit

Choose L4 for lower-power inference, media processing, and steady production serving where cost efficiency matters more than maximum training speed.

Cost decision checks for L4 buyers and renters.
Decision factor Buy or rent signal Cost check
Inference model size Choose L4 when 24GB VRAM is enough for the model, batch size, and latency target. If the model needs sharding across many L4s, compare against one larger GPU before committing.
Video and media workloads L4 is a good fit for transcoding, visual inference, and mixed AI/media pipelines. Include CPU, RAM, disk, and network charges because many providers bill the full instance shape.
Production serving Rent when demand is bursty; buy or reserve when utilization is steady and predictable. Model 24/7 monthly cost and committed-use discounts before comparing against hardware.
L4 vs T4 upgrade Move from T4 to L4 when the workload needs newer Tensor Core performance, AV1 media support, or more predictable production inference latency. Compare cost per successful request or video minute, not only the hourly price, because L4 can reduce replica count for some workloads.

L4 cost planning checklist

  • Start with the served model size and batch target. L4 is cost efficient when the workload fits comfortably in 24GB VRAM and does not need the memory bandwidth of larger training GPUs.
  • For managed cloud deployments, compare the complete machine price. CPU, memory, storage, network, and committed-use discounts can matter more than the visible per-GPU L4 rate.
  • For production inference, test cold start, concurrency, media codec support, and tail latency. The cheapest hourly rate is not useful if it forces extra replicas or slower response times.

Reference checks: Google Cloud GPU machine types · Google Cloud Run L4 pricing reference

Buy vs Rent Analysis

Buy Hardware
$2.8k
  • One-time cost, unlimited usage
  • Full control over hardware
  • Electricity & cooling costs extra
  • Depreciation over 2-3 years
Best if using >86420 hours total
Rent Cloud GPU
$0.032/hr
  • Pay only for what you use
  • No upfront investment
  • Scale up/down instantly
  • No maintenance required
Best for <86420 hours or variable usage
Breakeven Point
86,420
hours of usage

At $0.032/hr cloud pricing, buying the hardware pays off after 86,420 hours (~3601 days or 120.0 months of 24/7 usage).

Usage Monthly Cloud Cost Months to Breakeven
100 hrs/month $3.24 865 months
200 hrs/month $6.48 433 months
500 hrs/month $16.20 173 months

Cloud GPU Pricing

Rent L4 from 11 cloud providers. Prices shown per GPU per hour.

Provider Type Instance GPUs On-Demand Per GPU Spot Availability
Vast.ai marketplace vastai-l4 1x $0.032/hr $0.032/hr Cheapest $0.032/hr -
RunPod gpu-cloud NVIDIA L4 1x $0.490/hr $0.490/hr $0.490/hr -
Google Cloud Platform hyperscaler gcp-l4 1x $0.560/hr $0.560/hr $0.223/hr (-60%) -
Amazon Web Services hyperscaler g6e.xlarge 1x $1.86/hr $1.86/hr - -
Amazon Web Services hyperscaler g6e.2xlarge 1x $2.24/hr $2.24/hr - -
Amazon Web Services hyperscaler g6e.12xlarge 4x $10.49/hr $2.62/hr - -
Amazon Web Services hyperscaler g6e.4xlarge 1x $3.00/hr $3.00/hr - -
Amazon Web Services hyperscaler g6e.24xlarge 4x $15.07/hr $3.77/hr - -
Amazon Web Services hyperscaler g6e.48xlarge 8x $30.13/hr $3.77/hr - -
Amazon Web Services hyperscaler g6e.8xlarge 1x $4.53/hr $4.53/hr - -
Amazon Web Services hyperscaler g6e.16xlarge 1x $7.58/hr $7.58/hr - -
Best Spot Deal: Google Cloud Platform offers spot pricing at $0.223/hr (60% off on-demand).

L4 vs Alternatives

Compare L4 with similar GPUs from other brands.

GPU VRAM FP16 TFLOPS Bandwidth Hardware Price Cloud Price
L4 Current 22GB 60.6 300 GB/s $2.8k - -
AMD Radeon RX 7900 XTX AMD 24GB (+9%) 122.0 (+101%) 960 GB/s - - Compare
AMD Radeon RX 7900 XT AMD 20GB (-9%) 104.0 (+72%) 800 GB/s - - Compare
AMD Instinct MI100 AMD 32GB (+45%) 184.6 (+205%) 1.2 TB/s - - Compare

Best Use Cases

No specific use case recommendations for L4 yet.

Browse All Use Cases →

Compare L4

Other NVIDIA GPUs

Frequently Asked Questions about L4

L4 hardware and cloud prices vary by seller and region. This page compares the current site data, cloud options, and nearby alternatives so you can judge value instead of relying on one quote.

Yes, L4 is a strong inference GPU for smaller LLMs, embeddings, image workloads, and video pipelines when 24GB VRAM is enough.

NVIDIA L4 has 24GB of GPU memory. That makes it useful for many inference and media workloads, but large models may still require quantization, batching limits, or a larger GPU.

Compare the complete instance price, included CPU and memory, minimum billing unit, region, discounts, and measured requests or video minutes per hour. A lower GPU line item can still cost more when the required machine shape or idle capacity is larger.

L40S is a different, higher-performance 48GB GPU and should not be treated as another L4 listing. Use L40S only as an alternative when 24GB L4 memory or throughput is not enough.

L4 is usually the stronger choice for new inference and video workloads because it is newer, faster, and has better media acceleration. T4 can still be cheaper for legacy jobs that do not need the extra performance.

The L4 has a market price of approximately $2.8k (MSRP: $2.5k). Cloud rental starts at $0.032/hr. Prices may vary based on retailer, region, and availability.

Yes, the L4 with 22GB VRAM is suitable for many AI/ML workloads. For large language models, you may need multiple GPUs or consider higher-VRAM options like A100 or H100.

The breakeven point is approximately 86,420 hours of usage. Buy if you'll use it more than this; rent for shorter projects or variable workloads. Cloud rental from Vast.ai starts at $0.032/hr.

With 22GB VRAM and 60.6 FP16 TFLOPS, the L4 can run: Stable Diffusion, smaller LLMs (7B quantized), deep learning training, and gaming at high settings.

The L4 offers 22GB VRAM and 60.6 FP16 performance at $2.8k. Compare with similar GPUs using our comparison tool above. Key factors: VRAM for model size, TFLOPS for speed, and price for budget.
Sponsored

Rent GPUs from RunPod or Vast.ai

Estimate cost first