onrup

Gemma · Google

Fine-tuning Gemma 3 4B

What does it take to fine-tune Gemma 3 4B?

Gemma 3 4B needs 12 GB for half-precision LoRA training and 8 GB in four-bit, and 9 GB to serve. It supports supervised fine-tuning, under the Gemma Terms of Use licence. The cheapest qualifying class is RTX 3080 at $0.09 per GPU-hour.

The size at which Gemma 3 becomes generally useful rather than specialised. Instruction-tuned before you start, so a supervised fine-tune is adapting an existing conversational model rather than teaching one from a base.

What to know before choosing it

This is where Gemma 3 stops being specialised and becomes generally useful. Instruction-tuned already, so the fine-tune teaches your specifics on top of working conversational behaviour, which usually means less data than starting from a raw base would need.

Supervised fine-tuning only, across the whole Gemma family here. If your task needs preference tuning, this family is not a candidate however well it performs — worth checking at the start rather than after the dataset is built.

Specification

Gemma 3 4B specification
Parameters4B
ArchitectureDense
Size tierSmall
LoRA training memory12 GBHalf precision, frozen base
QLoRA training memory8 GBFour-bit base, higher-precision adapter
Serving memory9 GBHalf precision, before attention cache
ObjectivesSFT
LicenceGemma Terms of Use
Repositorygoogle/gemma-3-4b-it

What it costs to train

Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.

GPU classVRAMTraining / hrServing / hrFits
RTX 308012 GB$0.09$0.11LoRA and QLoRA
RTX 4000 Ada20 GB$0.09$0.11LoRA and QLoRA
L424 GB$0.17$0.20LoRA and QLoRA
RTX 309024 GB$0.22$0.26LoRA and QLoRA
RTX 409024 GB$0.38$0.43LoRA and QLoRA
A4048 GB$0.42$0.48LoRA and QLoRA
RTX 6000 Ada48 GB$0.61$0.71LoRA and QLoRA
A600048 GB$0.65$0.75LoRA and QLoRA
L40S48 GB$0.78$0.90LoRA and QLoRA
A100 40 GB40 GB$1.17$1.35LoRA and QLoRA
A100 80 GB80 GB$1.68$1.94LoRA and QLoRA
H100 80 GB80 GB$2.59$2.99LoRA and QLoRA
H200141 GB$4.55$5.25LoRA and QLoRA

Serving fits on RTX 3080 at $0.11 per hour — before the attention cache, which grows with context length and concurrency.

Good starting point for

Other sizes in this family

ModelParamsQLoRAObjectives
Gemma 3 1B1B3 GBSFT
Gemma 3 12B12B18 GBSFT
Gemma 3 27B27B36 GBSFT

Comparable sizes elsewhere

Frequently asked questions

How much VRAM does it take to fine-tune Gemma 3 4B?

12 GB for half-precision LoRA and 8 GB for four-bit QLoRA. Serving needs 9 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.

What is the cheapest way to fine-tune Gemma 3 4B?

Four-bit QLoRA on RTX 3080 at $0.09 per GPU-hour is the cheapest class that meets the 8 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.

Can I download the weights after fine-tuning Gemma 3 4B?

Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.

What licence does Gemma 3 4B carry?

Gemma Terms of Use. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.

Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.

Fine-tune Gemma 3 4B

Upload a dataset, forecast the run, and see the cost before any compute is leased.