onrup

GPU class · 48 GB

A40

What does A40 cost and what fits on it?

A40 has 48 GB and costs $0.42 per hour to train on, $0.48 to serve. 36 models in the catalogue fit half-precision LoRA on it and 5 more fit in four-bit.

Forty-eight gigabytes at close to 4090 money. That extra headroom is what makes preference tuning practical, because DPO holds a frozen reference model in memory alongside the one being trained and a 24GB card runs out well before a 14B pair fits.

When to pick this class

Forty-eight gigabytes at close to 4090 money is the whole argument, and the specific thing it unlocks is preference tuning. DPO holds a frozen reference model in memory alongside the one being trained, and a 24GB card runs out well before a 14B pair fits — so this is the cheapest class where preference tuning at mid-large scale is practical.

Against the RTX 6000 Ada at the same capacity it is slower and considerably cheaper. For a long run the faster card can win on total cost; for a short one this is the value pick and the difference is not close.

Rates

A40 rates
VRAM48 GB
Training$0.42 / GPU-hourMetered per GPU-second, one-minute floor
Serving$0.48 / GPU-hourWhile a replica is resident
Warm for a day$11.5224 hours resident, regardless of traffic
Warm for 30 days$345.60Set the endpoint to scale to zero if nobody is waiting
Models — LoRA36
Models — QLoRA only5

Largest models that fit for LoRA

Half precision, frozen base, within 48 GB.

Models that need four-bit training here

These exceed 48 GB in half precision but fit quantised. Quantised training is slower per step, so compare the total run cost against the next class up rather than the hourly rate alone.

Compared with its neighbours

The classes immediately either side of A40 on capacity and rate.

ClassVRAMTrainingServingVersus this one
RTX 409024 GB$0.38$0.4324 GB less, $0.04/hr cheaper
A100 40 GB40 GB$1.17$1.358 GB less, $0.75/hr dearer
RTX 6000 Ada48 GB$0.61$0.71same capacity, $0.19/hr dearer
A600048 GB$0.65$0.75same capacity, $0.23/hr dearer

The full rate card, all 13 classes →

Frequently asked questions

How much does A40 cost per hour?

$0.42 per GPU-hour for training and $0.48 for serving, metered per GPU-second above a one-minute floor. A continuously warm endpoint on this class is about $11.52 a day, or $345.60 over thirty days.

What models fit on A40?

36 models in the catalogue fit half-precision LoRA training within 48 GB, and 5 more fit in four-bit. 37 can be served from this class before accounting for the attention cache.

Is A40 the cheapest option?

Cheapest per hour is not the same as cheapest per run. Total cost is the rate multiplied by wall time, so a faster class that finishes sooner often costs less — particularly when the cheaper alternative would require quantised training, which is slower per step.

Last verified 6 August 2026.

Start with the free tier

A magic link creates your account, your tenant and your first API key. No card until you ask for compute.