GPU class · 24 GB
RTX 3090
What does RTX 3090 cost and what fits on it?
RTX 3090 has 24 GB and costs $0.22 per hour to train on, $0.26 to serve. 31 models in the catalogue fit half-precision LoRA on it and 3 more fit in four-bit.
The budget QLoRA workhorse. Twenty-four gigabytes at a mainstream price is the cheapest honest way to fine-tune a 7B or 8B model in four-bit, and it will do LoRA on anything up to about 4B in half precision. Slower than a 4090 on the same job, and correspondingly cheaper per hour.
When to pick this class
The budget QLoRA workhorse, and the cheapest honest way to fine-tune a 7B or 8B model in four-bit. Against the 4090 at the same capacity, it is slower per step and costs about forty per cent less per hour — so the total run cost is often close, and the 3090 wins when the run is short enough that throughput does not dominate.
Twenty-four gigabytes is also enough for half-precision LoRA on anything up to about 4B, which is a different and often better use of it: half-precision training on a small model beats quantised training on a larger one more often than people expect.
Rates
| VRAM | 24 GB |
|---|---|
| Training | $0.22 / GPU-hourMetered per GPU-second, one-minute floor |
| Serving | $0.26 / GPU-hourWhile a replica is resident |
| Warm for a day | $6.2424 hours resident, regardless of traffic |
| Warm for 30 days | $187.20Set the endpoint to scale to zero if nobody is waiting |
| Models — LoRA | 31 |
| Models — QLoRA only | 3 |
Largest models that fit for LoRA
Half precision, frozen base, within 24 GB.
Mistral Nemo 12B
12BNeeds 24 GB · Apache 2.0
Gemma 3 12B
12BNeeds 24 GB · Gemma Terms of Use
Falcon 3 10B
10BNeeds 22 GB · TII Falcon LLM
LFM2 8B-A1B
8.3BNeeds 18 GB · LFM Open
Qwen3 8B
8BNeeds 18 GB · Apache 2.0
Llama 3.1 8B
8BNeeds 18 GB · Llama 3.1 Community
DeepSeek-R1 Distill Llama 8B
8BNeeds 18 GB · Llama 3.1 Community
Granite 3.0 8B
8BNeeds 18 GB · Apache 2.0
Models that need four-bit training here
These exceed 24 GB in half precision but fit quantised. Quantised training is slower per step, so compare the total run cost against the next class up rather than the hourly rate alone.
Compared with its neighbours
The classes immediately either side of RTX 3090 on capacity and rate.
| Class | VRAM | Training | Serving | Versus this one |
|---|---|---|---|---|
| RTX 4000 Ada | 20 GB | $0.09 | $0.11 | 4 GB less, $0.13/hr cheaper |
| L4 | 24 GB | $0.17 | $0.20 | same capacity, $0.05/hr cheaper |
| RTX 4090 | 24 GB | $0.38 | $0.43 | same capacity, $0.16/hr dearer |
| A100 40 GB | 40 GB | $1.17 | $1.35 | 16 GB more, $0.95/hr dearer |
Frequently asked questions
How much does RTX 3090 cost per hour?
$0.22 per GPU-hour for training and $0.26 for serving, metered per GPU-second above a one-minute floor. A continuously warm endpoint on this class is about $6.24 a day, or $187.20 over thirty days.
What models fit on RTX 3090?
31 models in the catalogue fit half-precision LoRA training within 24 GB, and 3 more fit in four-bit. 32 can be served from this class before accounting for the attention cache.
Is RTX 3090 the cheapest option?
Cheapest per hour is not the same as cheapest per run. Total cost is the rate multiplied by wall time, so a faster class that finishes sooner often costs less — particularly when the cheaper alternative would require quantised training, which is slower per step.
Last verified 6 August 2026.
Start with the free tier
A magic link creates your account, your tenant and your first API key. No card until you ask for compute.