GPU class · 12 GB
RTX 3080
What does RTX 3080 cost and what fits on it?
RTX 3080 has 12 GB and costs $0.09 per hour to train on, $0.11 to serve. 17 models in the catalogue fit half-precision LoRA on it and 6 more fit in four-bit.
The cheapest class on the rate card and the right one for models at or below about two billion parameters. Twelve gigabytes is enough for LoRA on a 1B-class model with room for a reasonable batch size, and enough to serve one comfortably. It is not enough for anything at 7B, even quantised, once the optimiser state is accounted for.
When to pick this class
The reason to be on this class is that the model is genuinely small and the volume is genuinely large. At twelve gigabytes you are training 1B-class models and serving them, and the economics only make sense when the request count is high enough that the difference between nine cents and twenty-two cents an hour compounds into real money.
The trap is assuming a 7B will fit in four-bit. It will not once optimiser state and activations are accounted for, and the run will fail at a random-looking step rather than at the start. Check the model page's admission threshold before choosing this class rather than after.
Rates
| VRAM | 12 GB |
|---|---|
| Training | $0.09 / GPU-hourMetered per GPU-second, one-minute floor |
| Serving | $0.11 / GPU-hourWhile a replica is resident |
| Warm for a day | $2.6424 hours resident, regardless of traffic |
| Warm for 30 days | $79.20Set the endpoint to scale to zero if nobody is waiting |
| Models — LoRA | 17 |
| Models — QLoRA only | 6 |
Largest models that fit for LoRA
Half precision, frozen base, within 12 GB.
Qwen3 4B
4BNeeds 12 GB · Apache 2.0
Gemma 3 4B
4BNeeds 12 GB · Gemma Terms of Use
Phi-4 Mini 3.8B
3.8BNeeds 12 GB · MIT
SmolLM3 3B
3BNeeds 10 GB · Apache 2.0
Llama 3.2 3B
3BNeeds 10 GB · Llama 3.2 Community
Falcon 3 3B
3BNeeds 10 GB · TII Falcon LLM
LFM2 2.6B
2.6BNeeds 8 GB · LFM Open
Granite 3.0 2B
2BNeeds 7 GB · Apache 2.0
Models that need four-bit training here
These exceed 12 GB in half precision but fit quantised. Quantised training is slower per step, so compare the total run cost against the next class up rather than the hourly rate alone.
LFM2 8B-A1B
8.3B18 GB half precision · 12 GB in four-bit
Mistral 7B v0.3
7B16 GB half precision · 12 GB in four-bit
DeepSeek-R1 Distill Qwen 7B
7B16 GB half precision · 12 GB in four-bit
Falcon 3 7B
7B16 GB half precision · 12 GB in four-bit
OLMo 2 7B
7B16 GB half precision · 12 GB in four-bit
OLMo 3 7B
7B16 GB half precision · 12 GB in four-bit
Compared with its neighbours
The classes immediately either side of RTX 3080 on capacity and rate.
| Class | VRAM | Training | Serving | Versus this one |
|---|---|---|---|---|
| RTX 4000 Ada | 20 GB | $0.09 | $0.11 | 8 GB more, same rate |
| L4 | 24 GB | $0.17 | $0.20 | 12 GB more, $0.08/hr dearer |
Frequently asked questions
How much does RTX 3080 cost per hour?
$0.09 per GPU-hour for training and $0.11 for serving, metered per GPU-second above a one-minute floor. A continuously warm endpoint on this class is about $2.64 a day, or $79.20 over thirty days.
What models fit on RTX 3080?
17 models in the catalogue fit half-precision LoRA training within 12 GB, and 6 more fit in four-bit. 17 can be served from this class before accounting for the attention cache.
Is RTX 3080 the cheapest option?
Cheapest per hour is not the same as cheapest per run. Total cost is the rate multiplied by wall time, so a faster class that finishes sooner often costs less — particularly when the cheaper alternative would require quantised training, which is slower per step.
Last verified 6 August 2026.
Start with the free tier
A magic link creates your account, your tenant and your first API key. No card until you ask for compute.