DeepSeek-R1 distills · DeepSeek
Fine-tuning DeepSeek-R1 Distill Qwen 7B
What does it take to fine-tune DeepSeek-R1 Distill Qwen 7B?
DeepSeek-R1 Distill Qwen 7B needs 16 GB for half-precision LoRA training and 12 GB in four-bit, and 14 GB to serve. It supports supervised fine-tuning, preference tuning (dpo), under the Apache 2.0 licence. The cheapest qualifying class is RTX 3080 at $0.09 per GPU-hour.
A Qwen2.5 7B base with reasoning behaviour already distilled into it. For maths and code, this starts you several thousand training steps ahead of a plain base model, because you are refining a reasoning style rather than instilling one.
What to know before choosing it
Starting from a model that already reasons step by step is worth several thousand training steps on maths and code. You are refining an existing behaviour rather than instilling one, and the difference shows up as needing less data rather than as a higher ceiling.
The behaviour is also sticky in a way that matters: fine-tuning on terse answers will suppress the visible working, which is sometimes what you want and sometimes destroys the reason you chose this base. Decide which before building the dataset.
Specification
| Parameters | 7B |
|---|---|
| Architecture | Dense |
| Size tier | Mid |
| LoRA training memory | 16 GBHalf precision, frozen base |
| QLoRA training memory | 12 GBFour-bit base, higher-precision adapter |
| Serving memory | 14 GBHalf precision, before attention cache |
| Objectives | SFT, DPO |
| Licence | Apache 2.0 |
| Repository | deepseek-ai/DeepSeek-R1-Distill-Qwen-7B |
What it costs to train
Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.
| GPU class | VRAM | Training / hr | Serving / hr | Fits |
|---|---|---|---|---|
| RTX 3080 | 12 GB | $0.09 | $0.11 | QLoRA only |
| RTX 4000 Ada | 20 GB | $0.09 | $0.11 | LoRA and QLoRA |
| L4 | 24 GB | $0.17 | $0.20 | LoRA and QLoRA |
| RTX 3090 | 24 GB | $0.22 | $0.26 | LoRA and QLoRA |
| RTX 4090 | 24 GB | $0.38 | $0.43 | LoRA and QLoRA |
| A40 | 48 GB | $0.42 | $0.48 | LoRA and QLoRA |
| RTX 6000 Ada | 48 GB | $0.61 | $0.71 | LoRA and QLoRA |
| A6000 | 48 GB | $0.65 | $0.75 | LoRA and QLoRA |
| L40S | 48 GB | $0.78 | $0.90 | LoRA and QLoRA |
| A100 40 GB | 40 GB | $1.17 | $1.35 | LoRA and QLoRA |
| A100 80 GB | 80 GB | $1.68 | $1.94 | LoRA and QLoRA |
| H100 80 GB | 80 GB | $2.59 | $2.99 | LoRA and QLoRA |
| H200 | 141 GB | $4.55 | $5.25 | LoRA and QLoRA |
Half-precision LoRA needs 16 GB, so the cheapest class for it is RTX 4000 Ada at $0.09 per hour. Below that, training has to be quantised.
Serving fits on RTX 4000 Ada at $0.11 per hour — before the attention cache, which grows with context length and concurrency.
Good starting point for
- Mathematical reasoning
- Code generation with visible working
- Chain-of-thought tasks on a 24GB budget
Preference tuning on this model
Preference tuning holds a frozen reference copy of the model alongside the one being trained, so budget roughly 32 GB rather than 16 GB. That is the thing that catches people out — supervised training on this model fits on a class that preference tuning will overflow.
Other sizes in this family
| Model | Params | QLoRA | Objectives |
|---|---|---|---|
| DeepSeek-R1 Distill Llama 8B | 8B | 14 GB | SFT, DPO |
Comparable sizes elsewhere
Frequently asked questions
How much VRAM does it take to fine-tune DeepSeek-R1 Distill Qwen 7B?
16 GB for half-precision LoRA and 12 GB for four-bit QLoRA. Serving needs 14 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.
What is the cheapest way to fine-tune DeepSeek-R1 Distill Qwen 7B?
Four-bit QLoRA on RTX 3080 at $0.09 per GPU-hour is the cheapest class that meets the 12 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.
Can I download the weights after fine-tuning DeepSeek-R1 Distill Qwen 7B?
Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.
What licence does DeepSeek-R1 Distill Qwen 7B carry?
Apache 2.0. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.
Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.
Fine-tune DeepSeek-R1 Distill Qwen 7B
Upload a dataset, forecast the run, and see the cost before any compute is leased.