LFM2 · Liquid AI
Fine-tuning LFM2 700M
What does it take to fine-tune LFM2 700M?
LFM2 700M needs 4 GB for half-precision LoRA training and 3 GB in four-bit, and 2 GB to serve. It supports supervised fine-tuning, under the LFM Open licence. The cheapest qualifying class is RTX 3080 at $0.09 per GPU-hour.
Twice the parameters of the 350M and the same two gigabytes to serve, which makes it the better default of the pair when the deployment target is fixed and the only question is how much capability fits inside it.
What to know before choosing it
Twice the parameters of the 350M and the same two gigabytes to serve. When the deployment target is fixed — a device, a container with a hard memory limit — that makes this strictly the better of the pair, and the decision needs no further analysis.
Like the rest of the family it is supervised fine-tuning only, and the hybrid architecture means recipes written for standard transformers do not always transfer cleanly. Budget a little more time for the first run than you would for a Llama of the same size.
Specification
| Parameters | 700M |
|---|---|
| Architecture | Hybrid |
| Size tier | Tiny |
| LoRA training memory | 4 GBHalf precision, frozen base |
| QLoRA training memory | 3 GBFour-bit base, higher-precision adapter |
| Serving memory | 2 GBHalf precision, before attention cache |
| Objectives | SFT |
| Licence | LFM Open |
| Repository | LiquidAI/LFM2-700M |
What it costs to train
Every class with enough memory for four-bit training, cheapest first. Total cost is the rate multiplied by wall time, so the cheapest rate is not always the cheapest run — a faster class that finishes sooner frequently wins.
| GPU class | VRAM | Training / hr | Serving / hr | Fits |
|---|---|---|---|---|
| RTX 3080 | 12 GB | $0.09 | $0.11 | LoRA and QLoRA |
| RTX 4000 Ada | 20 GB | $0.09 | $0.11 | LoRA and QLoRA |
| L4 | 24 GB | $0.17 | $0.20 | LoRA and QLoRA |
| RTX 3090 | 24 GB | $0.22 | $0.26 | LoRA and QLoRA |
| RTX 4090 | 24 GB | $0.38 | $0.43 | LoRA and QLoRA |
| A40 | 48 GB | $0.42 | $0.48 | LoRA and QLoRA |
| RTX 6000 Ada | 48 GB | $0.61 | $0.71 | LoRA and QLoRA |
| A6000 | 48 GB | $0.65 | $0.75 | LoRA and QLoRA |
| L40S | 48 GB | $0.78 | $0.90 | LoRA and QLoRA |
| A100 40 GB | 40 GB | $1.17 | $1.35 | LoRA and QLoRA |
| A100 80 GB | 80 GB | $1.68 | $1.94 | LoRA and QLoRA |
| H100 80 GB | 80 GB | $2.59 | $2.99 | LoRA and QLoRA |
| H200 | 141 GB | $4.55 | $5.25 | LoRA and QLoRA |
Serving fits on RTX 3080 at $0.11 per hour — before the attention cache, which grows with context length and concurrency.
Good starting point for
- On-device assistants
- Mobile and embedded targets
- Latency-critical extraction
Other sizes in this family
| Model | Params | QLoRA | Objectives |
|---|---|---|---|
| LFM2 350M | 350M | 2 GB | SFT |
| LFM2 1.2B | 1.2B | 4 GB | SFT |
| LFM2 2.6B | 2.6B | 6 GB | SFT |
| LFM2 8B-A1B | 8.3B | 12 GB | SFT |
| LFM2 24B-A2B | 24B | 28 GB | SFT |
Comparable sizes elsewhere
Frequently asked questions
How much VRAM does it take to fine-tune LFM2 700M?
4 GB for half-precision LoRA and 3 GB for four-bit QLoRA. Serving needs 2 GB. Preference tuning roughly doubles the training figure, because a frozen reference model is held alongside the one being trained.
What is the cheapest way to fine-tune LFM2 700M?
Four-bit QLoRA on RTX 3080 at $0.09 per GPU-hour is the cheapest class that meets the 3 GB threshold. Note that quantised training is slower per step, so a faster class sometimes costs less over the whole run.
Can I download the weights after fine-tuning LFM2 700M?
Yes. Every finished run exposes its trained weights for download, and publishing to a model hub is a single call with a generated model card recording the base model and version the adapter applies to.
What licence does LFM2 700M carry?
LFM Open. The licence follows the fine-tune — a derivative inherits the base model’s terms, and those terms pass to anyone you give the model to.
Last verified 6 August 2026. Memory thresholds are the platform's own admission limits.
Fine-tune LFM2 700M
Upload a dataset, forecast the run, and see the cost before any compute is leased.