Gemma · Google
Fine-tuning Gemma 3 27B
What does it take to fine-tune Gemma 3 27B?
Gemma 3 27B needs 56 GB for half-precision LoRA training and 36 GB in four-bit, and 54 GB to serve. It supports supervised fine-tuning, under the Gemma Terms of Use licence. The cheapest qualifying class is A40 at $0.42 per GPU-hour.
The largest Gemma. Note that half-precision LoRA needs 56GB, which no 48GB class provides — so this is either a four-bit run on a 48GB card or a half-precision run on an 80GB one, and the price difference between those two paths is substantial.
What to know before choosing it
The memory numbers force a decision that is easy to miss: half-precision LoRA needs 56GB, which no 48GB class provides. So this is either a four-bit run on a 48GB card or a half-precision run on an 80GB one, and the cost difference between those paths is substantial.
It is the largest Gemma and the strongest multilingual generator in the catalogue at its size. Whether that is worth the licence terms and the supervised-only constraint is a question about your product rather than about the model.
Specification
| Parameters | 27B |
|---|---|
| Architecture | Dense |
| Size tier | Large |
| LoRA training memory | 56 GBHalf precision, frozen base |
| QLoRA training memory | 36 GBFour-bit base, higher-precision adapter |
| Serving memory | 54 GBHalf precision, before attention cache |
| Objectives | SFT |
| Licence | Gemma Terms of Use |
| Repository | google/gemma-3-27b-it |
What it costs to train
Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.
| GPU class | VRAM | Training / hr | Serving / hr | Fits |
|---|---|---|---|---|
| A40 | 48 GB | $0.42 | $0.48 | QLoRA only |
| RTX 6000 Ada | 48 GB | $0.61 | $0.71 | QLoRA only |
| A6000 | 48 GB | $0.65 | $0.75 | QLoRA only |
| L40S | 48 GB | $0.78 | $0.90 | QLoRA only |
| A100 40 GB | 40 GB | $1.17 | $1.35 | QLoRA only |
| A100 80 GB | 80 GB | $1.68 | $1.94 | LoRA and QLoRA |
| H100 80 GB | 80 GB | $2.59 | $2.99 | LoRA and QLoRA |
| H200 | 141 GB | $4.55 | $5.25 | LoRA and QLoRA |
Half-precision LoRA needs 56 GB, so the cheapest class for it is A100 80 GB at $1.68 per hour. Below that, training has to be quantised.
Serving fits on A100 80 GB at $1.94 per hour — before the attention cache, which grows with context length and concurrency.
Good starting point for
- Large-tier quality with Google-family behaviour
- Multilingual generation at scale
Other sizes in this family
| Model | Params | QLoRA | Objectives |
|---|---|---|---|
| Gemma 3 1B | 1B | 3 GB | SFT |
| Gemma 3 4B | 4B | 8 GB | SFT |
| Gemma 3 12B | 12B | 18 GB | SFT |
Comparable sizes elsewhere
Frequently asked questions
How much VRAM does it take to fine-tune Gemma 3 27B?
56 GB for half-precision LoRA and 36 GB for four-bit QLoRA. Serving needs 54 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.
What is the cheapest way to fine-tune Gemma 3 27B?
Four-bit QLoRA on A40 at $0.42 per GPU-hour is the cheapest class that meets the 36 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.
Can I download the weights after fine-tuning Gemma 3 27B?
Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.
What licence does Gemma 3 27B carry?
Gemma Terms of Use. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.
Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.
Fine-tune Gemma 3 27B
Upload a dataset, forecast the run, and see the cost before any compute is leased.