onrup

Qwen3 · Alibaba

Fine-tuning Qwen3 32B

What does it take to fine-tune Qwen3 32B?

Qwen3 32B needs 64 GB for half-precision LoRA training and 48 GB in four-bit, and 64 GB to serve. It supports supervised fine-tuning, under the Apache 2.0 licence. The cheapest qualifying class is A40 at $0.42 per GPU-hour.

Dense 32B under Apache-2.0, which is a rarer combination than it sounds. Four-bit training fits on a 48GB class; half precision needs 80GB. Supervised fine-tuning only — the memory cost of holding a reference model at this size puts preference objectives out of reach on a single GPU.

What to know before choosing it

Dense 32B under Apache-2.0 is a rarer combination than the size alone suggests, and it is the main argument for this model over anything else in the large tier. Most models at this scale carry community licences with terms that need reading.

Supervised fine-tuning only, and the reason is memory rather than policy: holding a frozen reference model alongside a 32B puts preference tuning past what a single GPU can do. If your project needs preference tuning at scale, the largest practical size is around 24B.

Specification

Qwen3 32B specification
Parameters32B
ArchitectureDense
Size tierLarge
LoRA training memory64 GBHalf precision, frozen base
QLoRA training memory48 GBFour-bit base, higher-precision adapter
Serving memory64 GBHalf precision, before attention cache
ObjectivesSFT
LicenceApache 2.0
RepositoryQwen/Qwen3-32B

What it costs to train

Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.

GPU classVRAMTraining / hrServing / hrFits
A4048 GB$0.42$0.48QLoRA only
RTX 6000 Ada48 GB$0.61$0.71QLoRA only
A600048 GB$0.65$0.75QLoRA only
L40S48 GB$0.78$0.90QLoRA only
A100 80 GB80 GB$1.68$1.94LoRA and QLoRA
H100 80 GB80 GB$2.59$2.99LoRA and QLoRA
H200141 GB$4.55$5.25LoRA and QLoRA

Half-precision LoRA needs 64 GB, so the cheapest class for it is A100 80 GB at $1.68 per hour. Below that, training has to be quantised.

Serving fits on A100 80 GB at $1.94 per hour — before the attention cache, which grows with context length and concurrency.

Good starting point for

Other sizes in this family

ModelParamsQLoRAObjectives
Qwen3 0.6B600M3 GBSFT, DPO, GRPO
Qwen3 1.7B1.7B4 GBSFT, DPO, GRPO
Qwen3 4B4B8 GBSFT, DPO, GRPO
Qwen3 8B8B14 GBSFT, DPO, GRPO
Qwen3 14B14B20 GBSFT, DPO, GRPO
Qwen3 30B-A3B30B36 GBSFT

Comparable sizes elsewhere

Frequently asked questions

How much VRAM does it take to fine-tune Qwen3 32B?

64 GB for half-precision LoRA and 48 GB for four-bit QLoRA. Serving needs 64 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.

What is the cheapest way to fine-tune Qwen3 32B?

Four-bit QLoRA on A40 at $0.42 per GPU-hour is the cheapest class that meets the 48 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.

Can I download the weights after fine-tuning Qwen3 32B?

Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.

What licence does Qwen3 32B carry?

Apache 2.0. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.

Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.

Fine-tune Qwen3 32B

Upload a dataset, forecast the run, and see the cost before any compute is leased.