Llama · Meta
Fine-tuning Llama 3.3 70B
What does it take to fine-tune Llama 3.3 70B?
Llama 3.3 70B needs 140 GB for half-precision LoRA training and 48 GB in four-bit, and 140 GB to serve. It supports supervised fine-tuning, under the Llama 3.3 Community licence. The cheapest qualifying class is A40 at $0.42 per GPU-hour.
Four-bit training only in practice: QLoRA fits within 48GB, but half-precision LoRA needs 140GB and serving needs the same again. The gap between what you can train and what you can afford to serve is wider here than anywhere else in the catalogue, and it is worth resolving that question before starting rather than after.
What to know before choosing it
The gap between what you can train and what you can afford to serve is wider here than anywhere else in the catalogue. Four-bit training fits inside a 48GB class; serving in half precision needs 140GB. Resolve that before starting, because a model you can train and cannot deploy is an expensive lesson.
In practice it means either quantised serving, with the quality question that raises, or accepting that this is a batch-generation model rather than an interactive one. Both are reasonable; neither is the default people assume.
Specification
| Parameters | 70B |
|---|---|
| Architecture | Dense |
| Size tier | Extra large |
| LoRA training memory | 140 GBHalf precision, frozen base |
| QLoRA training memory | 48 GBFour-bit base, higher-precision adapter |
| Serving memory | 140 GBHalf precision, before attention cache |
| Objectives | SFT |
| Licence | Llama 3.3 Community |
| Repository | meta-llama/Llama-3.3-70B-Instruct |
What it costs to train
Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.
| GPU class | VRAM | Training / hr | Serving / hr | Fits |
|---|---|---|---|---|
| A40 | 48 GB | $0.42 | $0.48 | QLoRA only |
| RTX 6000 Ada | 48 GB | $0.61 | $0.71 | QLoRA only |
| A6000 | 48 GB | $0.65 | $0.75 | QLoRA only |
| L40S | 48 GB | $0.78 | $0.90 | QLoRA only |
| A100 80 GB | 80 GB | $1.68 | $1.94 | QLoRA only |
| H100 80 GB | 80 GB | $2.59 | $2.99 | QLoRA only |
| H200 | 141 GB | $4.55 | $5.25 | LoRA and QLoRA |
Half-precision LoRA needs 140 GB, so the cheapest class for it is H200 at $4.55 per hour. Below that, training has to be quantised.
Serving fits on H200 at $5.25 per hour — before the attention cache, which grows with context length and concurrency.
Good starting point for
- Tasks where a 32B measurably falls short
- Batch generation rather than interactive serving
Other sizes in this family
| Model | Params | QLoRA | Objectives |
|---|---|---|---|
| Llama 3.2 1B | 1B | 3 GB | SFT, DPO |
| Llama 3.2 3B | 3B | 6 GB | SFT, DPO |
| Llama 3.1 8B | 8B | 14 GB | SFT, DPO, GRPO |
| Llama 4 Scout 109B | 109B | 80 GB | SFT |
Frequently asked questions
How much VRAM does it take to fine-tune Llama 3.3 70B?
140 GB for half-precision LoRA and 48 GB for four-bit QLoRA. Serving needs 140 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.
What is the cheapest way to fine-tune Llama 3.3 70B?
Four-bit QLoRA on A40 at $0.42 per GPU-hour is the cheapest class that meets the 48 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.
Can I download the weights after fine-tuning Llama 3.3 70B?
Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.
What licence does Llama 3.3 70B carry?
Llama 3.3 Community. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.
Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.
Fine-tune Llama 3.3 70B
Upload a dataset, forecast the run, and see the cost before any compute is leased.