GPU class · 24 GB
L4
What does L4 cost and what fits on it?
L4 has 24 GB and costs $0.17 per hour to train on, $0.20 to serve. 31 models in the catalogue fit half-precision LoRA on it and 3 more fit in four-bit.
A serving card rather than a training card. Twenty-four gigabytes in a low-power envelope makes it the economical choice for an endpoint that runs continuously at moderate traffic, where a 4090 would be faster per request but cost more per hour than the latency is worth.
When to pick this class
This is the class for an endpoint that runs continuously at moderate traffic. The comparison worth making is against the RTX 4090: the 4090 is faster per request and costs more than twice as much per hour, so the question is whether your users can tell. For batch work and background processing they cannot.
Its power envelope is also why availability tends to be good. Where a mainstream consumer card is competing for capacity against everyone else's training runs, this class is competing mostly against other serving workloads.
Rates
| VRAM | 24 GB |
|---|---|
| Training | $0.17 / GPU-hourMetered per GPU-second, one-minute floor |
| Serving | $0.20 / GPU-hourWhile a replica is resident |
| Warm for a day | $4.8024 hours resident, regardless of traffic |
| Warm for 30 days | $144Set the endpoint to scale to zero if nobody is waiting |
| Models — LoRA | 31 |
| Models — QLoRA only | 3 |
Largest models that fit for LoRA
Half precision, frozen base, within 24 GB.
Mistral Nemo 12B
12BNeeds 24 GB · Apache 2.0
Gemma 3 12B
12BNeeds 24 GB · Gemma Terms of Use
Falcon 3 10B
10BNeeds 22 GB · TII Falcon LLM
LFM2 8B-A1B
8.3BNeeds 18 GB · LFM Open
Qwen3 8B
8BNeeds 18 GB · Apache 2.0
Llama 3.1 8B
8BNeeds 18 GB · Llama 3.1 Community
DeepSeek-R1 Distill Llama 8B
8BNeeds 18 GB · Llama 3.1 Community
Granite 3.0 8B
8BNeeds 18 GB · Apache 2.0
Models that need four-bit training here
These exceed 24 GB in half precision but fit quantised. Quantised training is slower per step, so compare the total run cost against the next class up rather than the hourly rate alone.
Compared with its neighbours
The classes immediately either side of L4 on capacity and rate.
| Class | VRAM | Training | Serving | Versus this one |
|---|---|---|---|---|
| RTX 3080 | 12 GB | $0.09 | $0.11 | 12 GB less, $0.08/hr cheaper |
| RTX 4000 Ada | 20 GB | $0.09 | $0.11 | 4 GB less, $0.08/hr cheaper |
| RTX 3090 | 24 GB | $0.22 | $0.26 | same capacity, $0.05/hr dearer |
| RTX 4090 | 24 GB | $0.38 | $0.43 | same capacity, $0.21/hr dearer |
Frequently asked questions
How much does L4 cost per hour?
$0.17 per GPU-hour for training and $0.20 for serving, metered per GPU-second above a one-minute floor. A continuously warm endpoint on this class is about $4.80 a day, or $144 over thirty days.
What models fit on L4?
31 models in the catalogue fit half-precision LoRA training within 24 GB, and 3 more fit in four-bit. 32 can be served from this class before accounting for the attention cache.
Is L4 the cheapest option?
Cheapest per hour is not the same as cheapest per run. Total cost is the rate multiplied by wall time, so a faster class that finishes sooner often costs less — particularly when the cheaper alternative would require quantised training, which is slower per step.
Last verified 6 August 2026.
Start with the free tier
A magic link creates your account, your tenant and your first API key. No card until you ask for compute.