onrup

SmolLM · Hugging Face

Fine-tuning SmolLM3 3B

What does it take to fine-tune SmolLM3 3B?

SmolLM3 3B needs 10 GB for half-precision LoRA training and 6 GB in four-bit, and 7 GB to serve. It supports supervised fine-tuning, preference tuning (dpo), under the Apache 2.0 licence. The cheapest qualifying class is RTX 3080 at $0.09 per GPU-hour.

A dual-mode model that can answer directly or reason step by step, with six languages covered natively. At 3B it sits at the point where a small model stops being a toy for most structured tasks, and it still trains on a 24GB card without quantisation.

What to know before choosing it

The dual-mode behaviour is worth planning around rather than discovering. It can answer directly or work through a problem, and which mode it favours after fine-tuning is largely determined by what your training data looks like. A dataset of terse answers will suppress the reasoning mode almost entirely — usually what you want for extraction, usually not what you want for support.

Six languages natively is the other reason to pick it over a comparable 3B. Multilingual capability that survives fine-tuning is harder to get than it looks: a model that merely saw other languages in pretraining tends to lose them when you tune heavily on English.

Specification

SmolLM3 3B specification
Parameters3B
ArchitectureDense
Size tierSmall
LoRA training memory10 GBHalf precision, frozen base
QLoRA training memory6 GBFour-bit base, higher-precision adapter
Serving memory7 GBHalf precision, before attention cache
ObjectivesSFT, DPO
LicenceApache 2.0
RepositoryHuggingFaceTB/SmolLM3-3B

What it costs to train

Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.

GPU classVRAMTraining / hrServing / hrFits
RTX 308012 GB$0.09$0.11LoRA and QLoRA
RTX 4000 Ada20 GB$0.09$0.11LoRA and QLoRA
L424 GB$0.17$0.20LoRA and QLoRA
RTX 309024 GB$0.22$0.26LoRA and QLoRA
RTX 409024 GB$0.38$0.43LoRA and QLoRA
A4048 GB$0.42$0.48LoRA and QLoRA
RTX 6000 Ada48 GB$0.61$0.71LoRA and QLoRA
A600048 GB$0.65$0.75LoRA and QLoRA
L40S48 GB$0.78$0.90LoRA and QLoRA
A100 40 GB40 GB$1.17$1.35LoRA and QLoRA
A100 80 GB80 GB$1.68$1.94LoRA and QLoRA
H100 80 GB80 GB$2.59$2.99LoRA and QLoRA
H200141 GB$4.55$5.25LoRA and QLoRA

Serving fits on RTX 3080 at $0.11 per hour — before the attention cache, which grows with context length and concurrency.

Good starting point for

Preference tuning on this model

Preference tuning holds a frozen reference copy of the model alongside the one being trained, so budget roughly 20 GB rather than 10 GB. That is the thing that catches people out — supervised training on this model fits on a class that preference tuning will overflow.

When preference tuning beats supervised fine-tuning →

Other sizes in this family

ModelParamsQLoRAObjectives
SmolLM2 1.7B1.7B4 GBSFT, DPO

Comparable sizes elsewhere

Frequently asked questions

How much VRAM does it take to fine-tune SmolLM3 3B?

10 GB for half-precision LoRA and 6 GB for four-bit QLoRA. Serving needs 7 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.

What is the cheapest way to fine-tune SmolLM3 3B?

Four-bit QLoRA on RTX 3080 at $0.09 per GPU-hour is the cheapest class that meets the 6 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.

Can I download the weights after fine-tuning SmolLM3 3B?

Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.

What licence does SmolLM3 3B carry?

Apache 2.0. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.

Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.

Fine-tune SmolLM3 3B

Upload a dataset, forecast the run, and see the cost before any compute is leased.