Qwen3 · Alibaba
Fine-tuning Qwen3 14B
What does it take to fine-tune Qwen3 14B?
Qwen3 14B needs 28 GB for half-precision LoRA training and 20 GB in four-bit, and 26 GB to serve. It supports supervised fine-tuning, preference tuning (dpo), reinforcement learning (grpo), under the Apache 2.0 licence. The cheapest qualifying class is RTX 4000 Ada at $0.09 per GPU-hour.
The first rung where a 48GB class becomes necessary for half-precision LoRA. Worth the step from 8B when the task needs to hold more of the world in its head at once — long documents, dense domain knowledge, multi-hop reasoning — and not otherwise.
What to know before choosing it
The reason to step up from 8B is not general quality, which improves less than the parameter jump suggests. It is working capacity: holding a long document, a dense domain vocabulary and a multi-hop question at the same time. If your task does not need all three at once, 8B will usually match this at half the memory.
Twenty-eight gigabytes for half-precision LoRA is the awkward number — no 24GB class covers it, so this is where you either move to a 48GB class or accept quantised training. Compare the total run cost of both before assuming the cheaper card wins.
Specification
| Parameters | 14B |
|---|---|
| Architecture | Dense |
| Size tier | Mid-large |
| LoRA training memory | 28 GBHalf precision, frozen base |
| QLoRA training memory | 20 GBFour-bit base, higher-precision adapter |
| Serving memory | 26 GBHalf precision, before attention cache |
| Objectives | SFT, DPO, GRPO |
| Licence | Apache 2.0 |
| Repository | Qwen/Qwen3-14B |
What it costs to train
Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.
| GPU class | VRAM | Training / hr | Serving / hr | Fits |
|---|---|---|---|---|
| RTX 4000 Ada | 20 GB | $0.09 | $0.11 | QLoRA only |
| L4 | 24 GB | $0.17 | $0.20 | QLoRA only |
| RTX 3090 | 24 GB | $0.22 | $0.26 | QLoRA only |
| RTX 4090 | 24 GB | $0.38 | $0.43 | QLoRA only |
| A40 | 48 GB | $0.42 | $0.48 | LoRA and QLoRA |
| RTX 6000 Ada | 48 GB | $0.61 | $0.71 | LoRA and QLoRA |
| A6000 | 48 GB | $0.65 | $0.75 | LoRA and QLoRA |
| L40S | 48 GB | $0.78 | $0.90 | LoRA and QLoRA |
| A100 40 GB | 40 GB | $1.17 | $1.35 | LoRA and QLoRA |
| A100 80 GB | 80 GB | $1.68 | $1.94 | LoRA and QLoRA |
| H100 80 GB | 80 GB | $2.59 | $2.99 | LoRA and QLoRA |
| H200 | 141 GB | $4.55 | $5.25 | LoRA and QLoRA |
Half-precision LoRA needs 28 GB, so the cheapest class for it is A40 at $0.42 per hour. Below that, training has to be quantised.
Serving fits on A40 at $0.48 per hour — before the attention cache, which grows with context length and concurrency.
Good starting point for
- Long-document understanding
- Domain-heavy question answering
- Preference tuning where the reference model also has to fit
Preference tuning on this model
Preference tuning holds a frozen reference copy of the model alongside the one being trained, so budget roughly 56 GB rather than 28 GB. That is the thing that catches people out — supervised training on this model fits on a class that preference tuning will overflow.
Other sizes in this family
| Model | Params | QLoRA | Objectives |
|---|---|---|---|
| Qwen3 0.6B | 600M | 3 GB | SFT, DPO, GRPO |
| Qwen3 1.7B | 1.7B | 4 GB | SFT, DPO, GRPO |
| Qwen3 4B | 4B | 8 GB | SFT, DPO, GRPO |
| Qwen3 8B | 8B | 14 GB | SFT, DPO, GRPO |
| Qwen3 30B-A3B | 30B | 36 GB | SFT |
| Qwen3 32B | 32B | 48 GB | SFT |
Comparable sizes elsewhere
Frequently asked questions
How much VRAM does it take to fine-tune Qwen3 14B?
28 GB for half-precision LoRA and 20 GB for four-bit QLoRA. Serving needs 26 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.
What is the cheapest way to fine-tune Qwen3 14B?
Four-bit QLoRA on RTX 4000 Ada at $0.09 per GPU-hour is the cheapest class that meets the 20 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.
Can I download the weights after fine-tuning Qwen3 14B?
Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.
What licence does Qwen3 14B carry?
Apache 2.0. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.
Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.
Fine-tune Qwen3 14B
Upload a dataset, forecast the run, and see the cost before any compute is leased.