Mistral · Mistral AI
Fine-tuning Mistral Small 3 24B
What does it take to fine-tune Mistral Small 3 24B?
Mistral Small 3 24B needs 48 GB for half-precision LoRA training and 36 GB in four-bit, and 48 GB to serve. It supports supervised fine-tuning, preference tuning (dpo), under the Apache 2.0 licence. The cheapest qualifying class is A40 at $0.42 per GPU-hour.
Positioned by Mistral as best in class at 24B on release, and still Apache-2.0 at that size, which is unusual. It is the largest model in the catalogue that supports preference tuning, because above this the reference model stops fitting alongside the trained one.
What to know before choosing it
This is the largest model in the catalogue supporting preference tuning, and that is the specific reason to pick it. Above 24B the frozen reference model stops fitting alongside the trained one on a single GPU, so if your project needs DPO at the largest practical scale, this is where the ladder ends.
Apache-2.0 at 24B is also rare. Most models this size carry community terms, so for a commercial product where licence review is a real cost, this is often the best quality per unit of legal friction available.
Specification
| Parameters | 24B |
|---|---|
| Architecture | Dense |
| Size tier | Large |
| LoRA training memory | 48 GBHalf precision, frozen base |
| QLoRA training memory | 36 GBFour-bit base, higher-precision adapter |
| Serving memory | 48 GBHalf precision, before attention cache |
| Objectives | SFT, DPO |
| Licence | Apache 2.0 |
| Repository | mistralai/Mistral-Small-24B-Instruct-2501 |
What it costs to train
Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.
| GPU class | VRAM | Training / hr | Serving / hr | Fits |
|---|---|---|---|---|
| A40 | 48 GB | $0.42 | $0.48 | LoRA and QLoRA |
| RTX 6000 Ada | 48 GB | $0.61 | $0.71 | LoRA and QLoRA |
| A6000 | 48 GB | $0.65 | $0.75 | LoRA and QLoRA |
| L40S | 48 GB | $0.78 | $0.90 | LoRA and QLoRA |
| A100 40 GB | 40 GB | $1.17 | $1.35 | QLoRA only |
| A100 80 GB | 80 GB | $1.68 | $1.94 | LoRA and QLoRA |
| H100 80 GB | 80 GB | $2.59 | $2.99 | LoRA and QLoRA |
| H200 | 141 GB | $4.55 | $5.25 | LoRA and QLoRA |
Serving fits on A40 at $0.48 per hour — before the attention cache, which grows with context length and concurrency.
Good starting point for
- Preference tuning at the largest practical size
- High-quality generation under a permissive licence
Preference tuning on this model
Preference tuning holds a frozen reference copy of the model alongside the one being trained, so budget roughly 96 GB rather than 48 GB. That is the thing that catches people out — supervised training on this model fits on a class that preference tuning will overflow.
Other sizes in this family
| Model | Params | QLoRA | Objectives |
|---|---|---|---|
| Mistral 7B v0.3 | 7B | 12 GB | SFT, DPO, GRPO |
| Mistral Nemo 12B | 12B | 16 GB | SFT, DPO |
Comparable sizes elsewhere
Frequently asked questions
How much VRAM does it take to fine-tune Mistral Small 3 24B?
48 GB for half-precision LoRA and 36 GB for four-bit QLoRA. Serving needs 48 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.
What is the cheapest way to fine-tune Mistral Small 3 24B?
Four-bit QLoRA on A40 at $0.42 per GPU-hour is the cheapest class that meets the 36 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.
Can I download the weights after fine-tuning Mistral Small 3 24B?
Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.
What licence does Mistral Small 3 24B carry?
Apache 2.0. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.
Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.
Fine-tune Mistral Small 3 24B
Upload a dataset, forecast the run, and see the cost before any compute is leased.