onrup

GPU class · 48 GB

L40S

What does L40S cost and what fits on it?

L40S has 48 GB and costs $0.78 per hour to train on, $0.90 to serve. 36 models in the catalogue fit half-precision LoRA on it and 5 more fit in four-bit.

A newer datacentre 48GB part, and the class we reach for when an endpoint needs to hold many adapters against one base model at once. It is priced above the A40 and the RTX 6000 Ada, so for a single training run it is rarely the value pick; for a busy multi-tenant-style serving deployment inside one account, it usually is.

When to pick this class

This is the class we reach for when an endpoint needs to hold many adapters against one base model at once. For a single training run it is rarely the value pick — the A40 and the RTX 6000 Ada are both cheaper at the same capacity — but for a busy multi-variant serving deployment the throughput and memory bandwidth earn the difference.

If you are choosing it for training, check the cost calculator against the two cheaper 48GB classes first. If you are choosing it for serving several adapters under real traffic, it is usually correct and the alternatives are false economy.

Rates

L40S rates
VRAM48 GB
Training$0.78 / GPU-hourMetered per GPU-second, one-minute floor
Serving$0.90 / GPU-hourWhile a replica is resident
Warm for a day$21.6024 hours resident, regardless of traffic
Warm for 30 days$648Set the endpoint to scale to zero if nobody is waiting
Models — LoRA36
Models — QLoRA only5

Largest models that fit for LoRA

Half precision, frozen base, within 48 GB.

Models that need four-bit training here

These exceed 48 GB in half precision but fit quantised. Quantised training is slower per step, so compare the total run cost against the next class up rather than the hourly rate alone.

Compared with its neighbours

The classes immediately either side of L40S on capacity and rate.

ClassVRAMTrainingServingVersus this one
RTX 6000 Ada48 GB$0.61$0.71same capacity, $0.17/hr cheaper
A600048 GB$0.65$0.75same capacity, $0.13/hr cheaper
A100 80 GB80 GB$1.68$1.9432 GB more, $0.90/hr dearer
H100 80 GB80 GB$2.59$2.9932 GB more, $1.81/hr dearer

The full rate card, all 13 classes →

Frequently asked questions

How much does L40S cost per hour?

$0.78 per GPU-hour for training and $0.90 for serving, metered per GPU-second above a one-minute floor. A continuously warm endpoint on this class is about $21.60 a day, or $648 over thirty days.

What models fit on L40S?

36 models in the catalogue fit half-precision LoRA training within 48 GB, and 5 more fit in four-bit. 37 can be served from this class before accounting for the attention cache.

Is L40S the cheapest option?

Cheapest per hour is not the same as cheapest per run. Total cost is the rate multiplied by wall time, so a faster class that finishes sooner often costs less — particularly when the cheaper alternative would require quantised training, which is slower per step.

Last verified 6 August 2026.

Start with the free tier

A magic link creates your account, your tenant and your first API key. No card until you ask for compute.