Llama · Meta
Fine-tuning Llama 3.1 8B
What does it take to fine-tune Llama 3.1 8B?
Llama 3.1 8B needs 18 GB for half-precision LoRA training and 14 GB in four-bit, and 16 GB to serve. It supports supervised fine-tuning, preference tuning (dpo), reinforcement learning (grpo), under the Llama 3.1 Community licence. The cheapest qualifying class is RTX 4000 Ada at $0.09 per GPU-hour.
The single most fine-tuned open-weight model at this scale, which matters less for its raw quality than for everything around it: adapters, datasets, evaluation harnesses and troubleshooting threads all assume it. When something goes wrong, somebody has already written up the answer.
What to know before choosing it
The ecosystem argument reaches its strongest form here. More published adapters, more datasets already formatted for it, more evaluation harnesses assuming it, and more people who have hit whatever you are about to hit. For a first production fine-tune that is worth more than a benchmark point or two.
The trade is the licence. Llama 3.1 is a community licence with an acceptable-use policy and naming requirements attached, and those follow the derivative to whoever you give it to. If your product ships models to customers, read it before choosing — Qwen3 8B and Mistral 7B are the Apache-licensed alternatives at similar capability.
Specification
| Parameters | 8B |
|---|---|
| Architecture | Dense |
| Size tier | Mid |
| LoRA training memory | 18 GBHalf precision, frozen base |
| QLoRA training memory | 14 GBFour-bit base, higher-precision adapter |
| Serving memory | 16 GBHalf precision, before attention cache |
| Objectives | SFT, DPO, GRPO |
| Licence | Llama 3.1 Community |
| Repository | meta-llama/Llama-3.1-8B |
What it costs to train
Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.
| GPU class | VRAM | Training / hr | Serving / hr | Fits |
|---|---|---|---|---|
| RTX 4000 Ada | 20 GB | $0.09 | $0.11 | LoRA and QLoRA |
| L4 | 24 GB | $0.17 | $0.20 | LoRA and QLoRA |
| RTX 3090 | 24 GB | $0.22 | $0.26 | LoRA and QLoRA |
| RTX 4090 | 24 GB | $0.38 | $0.43 | LoRA and QLoRA |
| A40 | 48 GB | $0.42 | $0.48 | LoRA and QLoRA |
| RTX 6000 Ada | 48 GB | $0.61 | $0.71 | LoRA and QLoRA |
| A6000 | 48 GB | $0.65 | $0.75 | LoRA and QLoRA |
| L40S | 48 GB | $0.78 | $0.90 | LoRA and QLoRA |
| A100 40 GB | 40 GB | $1.17 | $1.35 | LoRA and QLoRA |
| A100 80 GB | 80 GB | $1.68 | $1.94 | LoRA and QLoRA |
| H100 80 GB | 80 GB | $2.59 | $2.99 | LoRA and QLoRA |
| H200 | 141 GB | $4.55 | $5.25 | LoRA and QLoRA |
Serving fits on RTX 4000 Ada at $0.11 per hour — before the attention cache, which grows with context length and concurrency.
Good starting point for
- Production fine-tunes with the deepest ecosystem support
- Teams that want a well-trodden path
- Reasoning work via reinforcement learning
Preference tuning on this model
Preference tuning holds a frozen reference copy of the model alongside the one being trained, so budget roughly 36 GB rather than 18 GB. That is the thing that catches people out — supervised training on this model fits on a class that preference tuning will overflow.
Other sizes in this family
| Model | Params | QLoRA | Objectives |
|---|---|---|---|
| Llama 3.2 1B | 1B | 3 GB | SFT, DPO |
| Llama 3.2 3B | 3B | 6 GB | SFT, DPO |
| Llama 3.3 70B | 70B | 48 GB | SFT |
| Llama 4 Scout 109B | 109B | 80 GB | SFT |
Comparable sizes elsewhere
Frequently asked questions
How much VRAM does it take to fine-tune Llama 3.1 8B?
18 GB for half-precision LoRA and 14 GB for four-bit QLoRA. Serving needs 16 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.
What is the cheapest way to fine-tune Llama 3.1 8B?
Four-bit QLoRA on RTX 4000 Ada at $0.09 per GPU-hour is the cheapest class that meets the 14 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.
Can I download the weights after fine-tuning Llama 3.1 8B?
Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.
What licence does Llama 3.1 8B carry?
Llama 3.1 Community. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.
Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.
Fine-tune Llama 3.1 8B
Upload a dataset, forecast the run, and see the cost before any compute is leased.