GPU class · 20 GB
RTX 4000 Ada
What does RTX 4000 Ada cost and what fits on it?
RTX 4000 Ada has 20 GB and costs $0.09 per hour to train on, $0.11 to serve. 28 models in the catalogue fit half-precision LoRA on it and 5 more fit in four-bit.
Priced identically to the RTX 3080 but with eight more gigabytes and a newer architecture, which makes it the better default of the two for serving rather than training. The extra headroom matters most when a small model needs a long context or a large KV cache.
When to pick this class
Priced identically to the RTX 3080 with eight more gigabytes and a newer architecture, which makes the choice between them straightforward: take this one unless availability pushes you elsewhere. The extra headroom is what lets a small model run a long context without the attention cache pushing it over.
It is a serving card in practice. The throughput advantage over the 3080 is modest for training, and if you are training something large enough to care about throughput you have already outgrown twenty gigabytes.
Rates
| VRAM | 20 GB |
|---|---|
| Training | $0.09 / GPU-hourMetered per GPU-second, one-minute floor |
| Serving | $0.11 / GPU-hourWhile a replica is resident |
| Warm for a day | $2.6424 hours resident, regardless of traffic |
| Warm for 30 days | $79.20Set the endpoint to scale to zero if nobody is waiting |
| Models — LoRA | 28 |
| Models — QLoRA only | 5 |
Largest models that fit for LoRA
Half precision, frozen base, within 20 GB.
LFM2 8B-A1B
8.3BNeeds 18 GB · LFM Open
Qwen3 8B
8BNeeds 18 GB · Apache 2.0
Llama 3.1 8B
8BNeeds 18 GB · Llama 3.1 Community
DeepSeek-R1 Distill Llama 8B
8BNeeds 18 GB · Llama 3.1 Community
Granite 3.0 8B
8BNeeds 18 GB · Apache 2.0
Granite 4.1 8B
8BNeeds 18 GB · Apache 2.0
Mistral 7B v0.3
7BNeeds 16 GB · Apache 2.0
DeepSeek-R1 Distill Qwen 7B
7BNeeds 16 GB · Apache 2.0
Models that need four-bit training here
These exceed 20 GB in half precision but fit quantised. Quantised training is slower per step, so compare the total run cost against the next class up rather than the hourly rate alone.
Compared with its neighbours
The classes immediately either side of RTX 4000 Ada on capacity and rate.
Frequently asked questions
How much does RTX 4000 Ada cost per hour?
$0.09 per GPU-hour for training and $0.11 for serving, metered per GPU-second above a one-minute floor. A continuously warm endpoint on this class is about $2.64 a day, or $79.20 over thirty days.
What models fit on RTX 4000 Ada?
28 models in the catalogue fit half-precision LoRA training within 20 GB, and 5 more fit in four-bit. 29 can be served from this class before accounting for the attention cache.
Is RTX 4000 Ada the cheapest option?
Cheapest per hour is not the same as cheapest per run. Total cost is the rate multiplied by wall time, so a faster class that finishes sooner often costs less — particularly when the cheaper alternative would require quantised training, which is slower per step.
Last verified 6 August 2026.
Start with the free tier
A magic link creates your account, your tenant and your first API key. No card until you ask for compute.