onrup

Qwen3 · Alibaba

Fine-tuning Qwen3 8B

What does it take to fine-tune Qwen3 8B?

Qwen3 8B needs 18 GB for half-precision LoRA training and 14 GB in four-bit, and 16 GB to serve. It supports supervised fine-tuning, preference tuning (dpo), reinforcement learning (grpo), under the Apache 2.0 licence. The cheapest qualifying class is RTX 4000 Ada at $0.09 per GPU-hour.

The Qwen team's own framing is that the 8B base performs around the level of the previous generation's 14B. It is Apache-2.0, supports every objective offered here, and fits half-precision LoRA on a single 24GB card — which is why it is the most common starting point on the platform for anyone without a specific reason to use Llama.

What to know before choosing it

The comparison the Qwen team draws — the 8B base performing around the level of the previous generation's 14B — is the practical reason this shows up as the default so often. You get mid-large behaviour at mid-tier memory, under a licence with no acceptable-use addendum, no naming requirement and no revenue threshold.

The other reason is objective coverage. It is one of the few models here supporting supervised, preference and reinforcement learning, so a project can start supervised and later add a reward function without changing base model — and changing base model mid-project costs more than people expect.

Specification

Qwen3 8B specification
Parameters8B
ArchitectureDense
Size tierMid
LoRA training memory18 GBHalf precision, frozen base
QLoRA training memory14 GBFour-bit base, higher-precision adapter
Serving memory16 GBHalf precision, before attention cache
ObjectivesSFT, DPO, GRPO
LicenceApache 2.0
RepositoryQwen/Qwen3-8B

What it costs to train

Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.

GPU classVRAMTraining / hrServing / hrFits
RTX 4000 Ada20 GB$0.09$0.11LoRA and QLoRA
L424 GB$0.17$0.20LoRA and QLoRA
RTX 309024 GB$0.22$0.26LoRA and QLoRA
RTX 409024 GB$0.38$0.43LoRA and QLoRA
A4048 GB$0.42$0.48LoRA and QLoRA
RTX 6000 Ada48 GB$0.61$0.71LoRA and QLoRA
A600048 GB$0.65$0.75LoRA and QLoRA
L40S48 GB$0.78$0.90LoRA and QLoRA
A100 40 GB40 GB$1.17$1.35LoRA and QLoRA
A100 80 GB80 GB$1.68$1.94LoRA and QLoRA
H100 80 GB80 GB$2.59$2.99LoRA and QLoRA
H200141 GB$4.55$5.25LoRA and QLoRA

Serving fits on RTX 4000 Ada at $0.11 per hour — before the attention cache, which grows with context length and concurrency.

Good starting point for

Preference tuning on this model

Preference tuning holds a frozen reference copy of the model alongside the one being trained, so budget roughly 36 GB rather than 18 GB. That is the thing that catches people out — supervised training on this model fits on a class that preference tuning will overflow.

When preference tuning beats supervised fine-tuning →

Other sizes in this family

ModelParamsQLoRAObjectives
Qwen3 0.6B600M3 GBSFT, DPO, GRPO
Qwen3 1.7B1.7B4 GBSFT, DPO, GRPO
Qwen3 4B4B8 GBSFT, DPO, GRPO
Qwen3 14B14B20 GBSFT, DPO, GRPO
Qwen3 30B-A3B30B36 GBSFT
Qwen3 32B32B48 GBSFT

Comparable sizes elsewhere

Frequently asked questions

How much VRAM does it take to fine-tune Qwen3 8B?

18 GB for half-precision LoRA and 14 GB for four-bit QLoRA. Serving needs 16 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.

What is the cheapest way to fine-tune Qwen3 8B?

Four-bit QLoRA on RTX 4000 Ada at $0.09 per GPU-hour is the cheapest class that meets the 14 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.

Can I download the weights after fine-tuning Qwen3 8B?

Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.

What licence does Qwen3 8B carry?

Apache 2.0. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.

Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.

Fine-tune Qwen3 8B

Upload a dataset, forecast the run, and see the cost before any compute is leased.