gpt-oss · OpenAI
Fine-tuning gpt-oss 20B
What does it take to fine-tune gpt-oss 20B?
gpt-oss 20B needs 40 GB for half-precision LoRA training and 24 GB in four-bit, and 24 GB to serve. It supports supervised fine-tuning, under the Apache 2.0 licence. The cheapest qualifying class is L4 at $0.17 per GPU-hour.
OpenAI's smaller open-weight release: twenty-one billion total parameters, 3.6 billion active, Apache-2.0. Twenty-four gigabytes to serve means a mainstream class can host it, which is a better ratio than any dense model of comparable capability. Experimental in this catalogue.
What to know before choosing it
Twenty-four gigabytes to serve a model with twenty-one billion total parameters is the headline: a mainstream class can host it, which is a better ratio than any dense model of comparable capability manages.
Training is where the sparsity stops helping. Forty gigabytes for half precision, and the usual mixture-of-experts difficulties apply. It is marked experimental in this catalogue for that reason rather than because of the model's quality.
Specification
| Parameters | 20.9B3.6B active per token |
|---|---|
| Architecture | Mixture of experts |
| Size tier | Large |
| LoRA training memory | 40 GBHalf precision, frozen base |
| QLoRA training memory | 24 GBFour-bit base, higher-precision adapter |
| Serving memory | 24 GBHalf precision, before attention cache |
| Objectives | SFT |
| Licence | Apache 2.0 |
| Repository | openai/gpt-oss-20b |
What it costs to train
Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.
| GPU class | VRAM | Training / hr | Serving / hr | Fits |
|---|---|---|---|---|
| L4 | 24 GB | $0.17 | $0.20 | QLoRA only |
| RTX 3090 | 24 GB | $0.22 | $0.26 | QLoRA only |
| RTX 4090 | 24 GB | $0.38 | $0.43 | QLoRA only |
| A40 | 48 GB | $0.42 | $0.48 | LoRA and QLoRA |
| RTX 6000 Ada | 48 GB | $0.61 | $0.71 | LoRA and QLoRA |
| A6000 | 48 GB | $0.65 | $0.75 | LoRA and QLoRA |
| L40S | 48 GB | $0.78 | $0.90 | LoRA and QLoRA |
| A100 40 GB | 40 GB | $1.17 | $1.35 | LoRA and QLoRA |
| A100 80 GB | 80 GB | $1.68 | $1.94 | LoRA and QLoRA |
| H100 80 GB | 80 GB | $2.59 | $2.99 | LoRA and QLoRA |
| H200 | 141 GB | $4.55 | $5.25 | LoRA and QLoRA |
Half-precision LoRA needs 40 GB, so the cheapest class for it is A40 at $0.42 per hour. Below that, training has to be quantised.
Serving fits on L4 at $0.20 per hour — before the attention cache, which grows with context length and concurrency.
Good starting point for
- Serving large-model behaviour on mainstream hardware
- Tasks where OpenAI-family behaviour is the baseline
This is a mixture-of-experts model
20.9B total parameters with 3.6B active per token. Memory scales with the total; inference compute scales with the active count. That makes it attractive when memory is cheaper than compute for your workload, and unattractive when the reverse holds. Quoting only one of the two numbers is how these models get misrepresented in both directions.
Other sizes in this family
| Model | Params | QLoRA | Objectives |
|---|---|---|---|
| gpt-oss 120B | 116.8B | 80 GB | SFT |
Comparable sizes elsewhere
Frequently asked questions
How much VRAM does it take to fine-tune gpt-oss 20B?
40 GB for half-precision LoRA and 24 GB for four-bit QLoRA. Serving needs 24 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.
What is the cheapest way to fine-tune gpt-oss 20B?
Four-bit QLoRA on L4 at $0.17 per GPU-hour is the cheapest class that meets the 24 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.
Can I download the weights after fine-tuning gpt-oss 20B?
Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.
What licence does gpt-oss 20B carry?
Apache 2.0. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.
Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.
Fine-tune gpt-oss 20B
Upload a dataset, forecast the run, and see the cost before any compute is leased.