onrup

LFM2 · Liquid AI

Fine-tuning LFM2 350M

What does it take to fine-tune LFM2 350M?

LFM2 350M needs 3 GB for half-precision LoRA training and 2 GB in four-bit, and 2 GB to serve. It supports supervised fine-tuning, under the LFM Open licence. The cheapest qualifying class is RTX 3080 at $0.09 per GPU-hour.

The smallest model in the catalogue by an order of magnitude, using a hybrid of gated convolutions and grouped-query attention rather than pure attention. Two gigabytes to serve. Licence is Apache-2.0-derived but adds a commercial revenue threshold, so check it if you are past ten million in revenue.

What to know before choosing it

Two gigabytes to serve puts this in a different category from everything else in the catalogue — it is deployable where a GPU may not exist at all. The hybrid architecture, combining gated convolutions with attention, is what makes that footprint achievable at this quality.

The licence is Apache-2.0-derived but adds a commercial revenue threshold. That is fine for most teams and a blocker for some, and it is exactly the kind of clause that is cheaper to read now than to discover during a funding round.

Specification

LFM2 350M specification
Parameters350M
ArchitectureHybrid
Size tierTiny
LoRA training memory3 GBHalf precision, frozen base
QLoRA training memory2 GBFour-bit base, higher-precision adapter
Serving memory2 GBHalf precision, before attention cache
ObjectivesSFT
LicenceLFM Open
RepositoryLiquidAI/LFM2-350M

What it costs to train

Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.

GPU classVRAMTraining / hrServing / hrFits
RTX 308012 GB$0.09$0.11LoRA and QLoRA
RTX 4000 Ada20 GB$0.09$0.11LoRA and QLoRA
L424 GB$0.17$0.20LoRA and QLoRA
RTX 309024 GB$0.22$0.26LoRA and QLoRA
RTX 409024 GB$0.38$0.43LoRA and QLoRA
A4048 GB$0.42$0.48LoRA and QLoRA
RTX 6000 Ada48 GB$0.61$0.71LoRA and QLoRA
A600048 GB$0.65$0.75LoRA and QLoRA
L40S48 GB$0.78$0.90LoRA and QLoRA
A100 40 GB40 GB$1.17$1.35LoRA and QLoRA
A100 80 GB80 GB$1.68$1.94LoRA and QLoRA
H100 80 GB80 GB$2.59$2.99LoRA and QLoRA
H200141 GB$4.55$5.25LoRA and QLoRA

Serving fits on RTX 3080 at $0.11 per hour — before the attention cache, which grows with context length and concurrency.

Good starting point for

Other sizes in this family

ModelParamsQLoRAObjectives
LFM2 700M700M3 GBSFT
LFM2 1.2B1.2B4 GBSFT
LFM2 2.6B2.6B6 GBSFT
LFM2 8B-A1B8.3B12 GBSFT
LFM2 24B-A2B24B28 GBSFT

Comparable sizes elsewhere

Frequently asked questions

How much VRAM does it take to fine-tune LFM2 350M?

3 GB for half-precision LoRA and 2 GB for four-bit QLoRA. Serving needs 2 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.

What is the cheapest way to fine-tune LFM2 350M?

Four-bit QLoRA on RTX 3080 at $0.09 per GPU-hour is the cheapest class that meets the 2 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.

Can I download the weights after fine-tuning LFM2 350M?

Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.

What licence does LFM2 350M carry?

LFM Open. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.

Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.

Fine-tune LFM2 350M

Upload a dataset, forecast the run, and see the cost before any compute is leased.