Mistral · Mistral AI
Fine-tuning Mistral Nemo 12B
What does it take to fine-tune Mistral Nemo 12B?
Mistral Nemo 12B needs 24 GB for half-precision LoRA training and 16 GB in four-bit, and 24 GB to serve. It supports supervised fine-tuning, preference tuning (dpo), under the Apache 2.0 licence. The cheapest qualifying class is RTX 4000 Ada at $0.09 per GPU-hour.
A 128K-context model built with NVIDIA, still Apache-2.0. The long context is the reason to choose it: at 12B it is the cheapest model here that can hold an entire long document in a single pass without retrieval, which removes a whole layer of system complexity.
What to know before choosing it
The 128K context is the reason to choose this, and it removes a layer of system complexity rather than merely improving a number. A pipeline that can put a forty-page report in one pass needs no chunking, no merge logic, and does not have the seam errors chunking introduces.
Apache-2.0 at 12B with that context length is an unusual combination. The cost is that twenty-four gigabytes for half-precision LoRA sits exactly at the ceiling of the mainstream classes — it fits with nothing spare, so long sequences will push you to 48GB.
Specification
| Parameters | 12B |
|---|---|
| Architecture | Dense |
| Size tier | Mid-large |
| LoRA training memory | 24 GBHalf precision, frozen base |
| QLoRA training memory | 16 GBFour-bit base, higher-precision adapter |
| Serving memory | 24 GBHalf precision, before attention cache |
| Objectives | SFT, DPO |
| Licence | Apache 2.0 |
| Repository | mistralai/Mistral-Nemo-Base-2407 |
What it costs to train
Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.
| GPU class | VRAM | Training / hr | Serving / hr | Fits |
|---|---|---|---|---|
| RTX 4000 Ada | 20 GB | $0.09 | $0.11 | QLoRA only |
| L4 | 24 GB | $0.17 | $0.20 | LoRA and QLoRA |
| RTX 3090 | 24 GB | $0.22 | $0.26 | LoRA and QLoRA |
| RTX 4090 | 24 GB | $0.38 | $0.43 | LoRA and QLoRA |
| A40 | 48 GB | $0.42 | $0.48 | LoRA and QLoRA |
| RTX 6000 Ada | 48 GB | $0.61 | $0.71 | LoRA and QLoRA |
| A6000 | 48 GB | $0.65 | $0.75 | LoRA and QLoRA |
| L40S | 48 GB | $0.78 | $0.90 | LoRA and QLoRA |
| A100 40 GB | 40 GB | $1.17 | $1.35 | LoRA and QLoRA |
| A100 80 GB | 80 GB | $1.68 | $1.94 | LoRA and QLoRA |
| H100 80 GB | 80 GB | $2.59 | $2.99 | LoRA and QLoRA |
| H200 | 141 GB | $4.55 | $5.25 | LoRA and QLoRA |
Half-precision LoRA needs 24 GB, so the cheapest class for it is L4 at $0.17 per hour. Below that, training has to be quantised.
Serving fits on L4 at $0.20 per hour — before the attention cache, which grows with context length and concurrency.
Good starting point for
- Whole-document processing without retrieval
- Long conversation histories
- Contract and report analysis
Preference tuning on this model
Preference tuning holds a frozen reference copy of the model alongside the one being trained, so budget roughly 48 GB rather than 24 GB. That is the thing that catches people out — supervised training on this model fits on a class that preference tuning will overflow.
Other sizes in this family
| Model | Params | QLoRA | Objectives |
|---|---|---|---|
| Mistral 7B v0.3 | 7B | 12 GB | SFT, DPO, GRPO |
| Mistral Small 3 24B | 24B | 36 GB | SFT, DPO |
Comparable sizes elsewhere
Frequently asked questions
How much VRAM does it take to fine-tune Mistral Nemo 12B?
24 GB for half-precision LoRA and 16 GB for four-bit QLoRA. Serving needs 24 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.
What is the cheapest way to fine-tune Mistral Nemo 12B?
Four-bit QLoRA on RTX 4000 Ada at $0.09 per GPU-hour is the cheapest class that meets the 16 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.
Can I download the weights after fine-tuning Mistral Nemo 12B?
Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.
What licence does Mistral Nemo 12B carry?
Apache 2.0. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.
Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.
Fine-tune Mistral Nemo 12B
Upload a dataset, forecast the run, and see the cost before any compute is leased.