onrup

LFM2 · Liquid AI

Fine-tuning LFM2 24B-A2B

What does it take to fine-tune LFM2 24B-A2B?

LFM2 24B-A2B needs 48 GB for half-precision LoRA training and 28 GB in four-bit, and 32 GB to serve. It supports supervised fine-tuning, under the LFM Open licence. The cheapest qualifying class is A40 at $0.42 per GPU-hour.

The largest LFM2, with 2.3 billion active parameters out of twenty-four billion. Thirty-two gigabytes to serve a 24B model is unusually low, and it is the strongest argument in the catalogue for the mixture-of-experts trade when serving cost dominates.

What to know before choosing it

Thirty-two gigabytes to serve a 24B model is unusually low, and it is the strongest argument in the catalogue for the mixture-of-experts trade when serving cost dominates. Large-tier knowledge, mid-tier serving footprint.

Training is the harder half. Forty-eight gigabytes for half precision, and the usual mixture-of-experts caveat applies: routing can become unbalanced during training in a way that looks like the model simply plateauing.

Specification

LFM2 24B-A2B specification
Parameters24B2.3B active per token
ArchitectureMixture of experts
Size tierLarge
LoRA training memory48 GBHalf precision, frozen base
QLoRA training memory28 GBFour-bit base, higher-precision adapter
Serving memory32 GBHalf precision, before attention cache
ObjectivesSFT
LicenceLFM Open
RepositoryLiquidAI/LFM2-24B-A2B

What it costs to train

Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.

GPU classVRAMTraining / hrServing / hrFits
A4048 GB$0.42$0.48LoRA and QLoRA
RTX 6000 Ada48 GB$0.61$0.71LoRA and QLoRA
A600048 GB$0.65$0.75LoRA and QLoRA
L40S48 GB$0.78$0.90LoRA and QLoRA
A100 40 GB40 GB$1.17$1.35QLoRA only
A100 80 GB80 GB$1.68$1.94LoRA and QLoRA
H100 80 GB80 GB$2.59$2.99LoRA and QLoRA
H200141 GB$4.55$5.25LoRA and QLoRA

Serving fits on A40 at $0.48 per hour — before the attention cache, which grows with context length and concurrency.

Good starting point for

This is a mixture-of-experts model

24B total parameters with 2.3B active per token. Memory scales with the total; inference compute scales with the active count. That makes it attractive when memory is cheaper than compute for your workload, and unattractive when the reverse holds. Quoting only one of the two numbers is how these models get misrepresented in both directions.

What mixture of experts means →

Other sizes in this family

ModelParamsQLoRAObjectives
LFM2 350M350M2 GBSFT
LFM2 700M700M3 GBSFT
LFM2 1.2B1.2B4 GBSFT
LFM2 2.6B2.6B6 GBSFT
LFM2 8B-A1B8.3B12 GBSFT

Comparable sizes elsewhere

Frequently asked questions

How much VRAM does it take to fine-tune LFM2 24B-A2B?

48 GB for half-precision LoRA and 28 GB for four-bit QLoRA. Serving needs 32 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.

What is the cheapest way to fine-tune LFM2 24B-A2B?

Four-bit QLoRA on A40 at $0.42 per GPU-hour is the cheapest class that meets the 28 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.

Can I download the weights after fine-tuning LFM2 24B-A2B?

Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.

What licence does LFM2 24B-A2B carry?

LFM Open. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.

Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.

Fine-tune LFM2 24B-A2B

Upload a dataset, forecast the run, and see the cost before any compute is leased.