GPU class · 40 GB
A100 40 GB
What does A100 40 GB cost and what fits on it?
A100 40 GB has 40 GB and costs $1.17 per hour to train on, $1.35 to serve. 34 models in the catalogue fit half-precision LoRA on it and 4 more fit in four-bit.
The datacentre staple. Less VRAM than the 48GB consumer-derived classes but far more memory bandwidth, which is what actually limits training throughput on larger batches. Choose it over an A40 when the run is bandwidth-bound rather than capacity-bound. No fp8 support.
When to pick this class
Less VRAM than the 48GB consumer-derived classes and far more memory bandwidth, which is the property that actually limits training throughput once batches get large. Choose it over an A40 when the run is bandwidth-bound rather than capacity-bound — and if you are not sure which, it is capacity-bound.
No fp8 support, which matters only for serving and only against an H100. For training at this scale it is not a consideration.
Rates
| VRAM | 40 GB |
|---|---|
| Training | $1.17 / GPU-hourMetered per GPU-second, one-minute floor |
| Serving | $1.35 / GPU-hourWhile a replica is resident |
| Warm for a day | $32.4024 hours resident, regardless of traffic |
| Warm for 30 days | $972Set the endpoint to scale to zero if nobody is waiting |
| Models — LoRA | 34 |
| Models — QLoRA only | 4 |
Largest models that fit for LoRA
Half precision, frozen base, within 40 GB.
gpt-oss 20B
20.9BNeeds 40 GB · Apache 2.0
Qwen3 14B
14BNeeds 28 GB · Apache 2.0
Phi-4 14B
14BNeeds 28 GB · MIT
Mistral Nemo 12B
12BNeeds 24 GB · Apache 2.0
Gemma 3 12B
12BNeeds 24 GB · Gemma Terms of Use
Falcon 3 10B
10BNeeds 22 GB · TII Falcon LLM
LFM2 8B-A1B
8.3BNeeds 18 GB · LFM Open
Qwen3 8B
8BNeeds 18 GB · Apache 2.0
Models that need four-bit training here
These exceed 40 GB in half precision but fit quantised. Quantised training is slower per step, so compare the total run cost against the next class up rather than the hourly rate alone.
Compared with its neighbours
The classes immediately either side of A100 40 GB on capacity and rate.
| Class | VRAM | Training | Serving | Versus this one |
|---|---|---|---|---|
| RTX 3090 | 24 GB | $0.22 | $0.26 | 16 GB less, $0.95/hr cheaper |
| RTX 4090 | 24 GB | $0.38 | $0.43 | 16 GB less, $0.79/hr cheaper |
| A40 | 48 GB | $0.42 | $0.48 | 8 GB more, $0.75/hr cheaper |
| RTX 6000 Ada | 48 GB | $0.61 | $0.71 | 8 GB more, $0.56/hr cheaper |
Frequently asked questions
How much does A100 40 GB cost per hour?
$1.17 per GPU-hour for training and $1.35 for serving, metered per GPU-second above a one-minute floor. A continuously warm endpoint on this class is about $32.40 a day, or $972 over thirty days.
What models fit on A100 40 GB?
34 models in the catalogue fit half-precision LoRA training within 40 GB, and 4 more fit in four-bit. 35 can be served from this class before accounting for the attention cache.
Is A100 40 GB the cheapest option?
Cheapest per hour is not the same as cheapest per run. Total cost is the rate multiplied by wall time, so a faster class that finishes sooner often costs less — particularly when the cheaper alternative would require quantised training, which is slower per step.
Last verified 6 August 2026.
Start with the free tier
A magic link creates your account, your tenant and your first API key. No card until you ask for compute.