onrup

Market map

The fine-tuning and inference landscape

Who is in this market?

14 managed platforms across four categories: managed fine-tuning products, serverless GPU platforms, hyperscaler ML suites and closed model APIs. Every figure below comes from the vendor’s own pricing page on the date checked. Where a vendor publishes no comparable rate, the cell is blank rather than estimated.

We build a competing product, so read this with that in mind. What we can offer instead of neutrality is verifiability: every claim links to the vendor’s own page, and where our own product is the weaker option we say so on the comparison page for that vendor.

Raw GPU marketplaces are deliberately out of scope. They solve a different problem — renting hardware rather than running a fine-tuning workflow — and mixing the two categories makes every comparison misleading.

The whole field

PlatformCategoryTrainsWeights portableEval gateH100 / hr
OnrupManaged fine-tuning platformYesYesBlocking$2.99
Together AIManaged fine-tuning platformYesYes$5.49
Fireworks AIManaged fine-tuning platformYesPartial$7
ModalServerless GPU platformNoYes$3.95
BasetenServerless GPU platformYesYes$6.50
ReplicateServerless GPU platformYesPartial$5.49
OpenAI fine-tuningClosed model APIYesNo
Hugging FaceManaged fine-tuning platformYesYes
PredibaseManaged fine-tuning platformYesYes
OpenPipeManaged fine-tuning platformYesYes
AWS SageMakerHyperscaler ML platformYesYes
Google Vertex AIHyperscaler ML platformYesPartial
AWS BedrockClosed model APIYesNo
AnyscaleServerless GPU platformYesYes
LaminiManaged fine-tuning platformYesYesYes

A dash means the vendor does not publish a figure comparable to the others in that column, not that the number is unknown to us and guessed at.

Published H100 rates

5 of 14 publish a directly comparable on-demand H100 80GB rate.

PlatformPer GPU-hourSource, checked
Onrup$2.992026-08-06
Modal$3.952026-08-06
Together AI$5.492026-08-06
Replicate$5.492026-08-06
Baseten$6.502026-08-06
Fireworks AI$72026-08-06

Published per-token training rates

LoRA supervised fine-tuning, smallest model tier, per million training tokens. This is a different pricing shape from GPU-hours and the two do not convert cleanly — which is cheaper depends on your dataset size and sequence length. We do not publish a per-token rate, so we are absent from this table rather than estimated into it.

PlatformPer 1M training tokensSource, checked
Together AI$0.482026-08-06
Fireworks AI$0.502026-08-06

Managed fine-tuning platform

A broad model API with fine-tuning attached, covering the widest catalogue of open-weight models of anyone in this set.

Pricing. LoRA supervised fine-tuning is $0.48 per 1M tokens for models up to 16B, rising to $1.50 for 17–69B and $2.90 for 70–100B. Full fine-tuning is roughly 2.5× the LoRA rate at every tier. Dedicated H100 capacity is $5.49 per GPU-hour on demand.

Strongest at. Model breadth. If the specific base model you need is unusual, Together is more likely to have it than anyone else here, and their token-priced inference means you never think about a GPU at all.

Inference, fine-tuning and evaluation behind one API, with fine-tuned adapters served at the same per-token rate as the base model.

Pricing. LoRA supervised fine-tuning is $0.50 per 1M training tokens up to 16B, $3.00 for 16–80B and $6.00 for 80–300B; preference tuning is double the supervised rate at every tier. On-demand H100 and H200 are both $7.00 per GPU-hour.

Strongest at. Serving a fine-tuned adapter at base-model token rates is a genuinely good deal, and it means a low-traffic fine-tune costs nothing to keep available. If your traffic is spiky and light, per-token beats per-hour and they are the cleanest expression of that model.

The model hub itself, plus AutoTrain for training and Inference Endpoints for serving. The centre of gravity of the open-weight world.

Pricing. Inference Endpoints are billed per hour by instance type across several cloud providers, so there is no single H100 rate to quote. AutoTrain is billed as underlying compute rather than per training token.

Strongest at. Gravity. The models, the datasets, the leaderboards and the community are all there already, and publishing to the Hub is where a fine-tune becomes visible to anyone else. We publish to the Hub too, which should tell you how we rate it.

A fine-tuning platform built around serving many LoRA adapters from a single GPU. Acquired — predibase.com now redirects to Rubrik.

Pricing. No standalone pricing page is reachable: predibase.com returns a 301 redirect to rubrik.com as of the date checked, following the acquisition. Existing customers should confirm continuity terms directly.

Strongest at. Their multi-adapter serving work was the best in the category, and the open-source server that came out of it is still worth reading if you are serving hundreds of adapters against one base model.

Capture production traffic from a large model, then train a small one to replace it on that exact distribution.

Pricing. Priced per training and inference token by model rather than by GPU-hour; no comparable H100 hourly rate is published.

Strongest at. The capture-then-train workflow is the best answer in this set to "I have a working prompt and a big bill". If your training data is really your production logs, they have removed more of that specific work than anyone.

Enterprise fine-tuning with a focus on factual accuracy and hallucination reduction, deployable on-premises.

Pricing. Enterprise pricing by arrangement; no public rate card is published.

Strongest at. On-premises and air-gapped deployment, and specific technique work aimed at making a model stop inventing facts. If the model has to run inside your own building, most of this comparison set is simply unavailable to you.

Serverless GPU platform

Serverless compute for arbitrary Python, billed by the second. Not a fine-tuning product — a substrate you build one on.

Pricing. Priced per second of GPU time: H100 at $0.001097/sec, which is $3.95 per GPU-hour. A100 80GB is $2.50 and L40S $1.95 per hour. There is no fine-tuning product and therefore no per-token training rate; you write the training code yourself.

Strongest at. Flexibility, and it is not close. If your workload is not a standard fine-tune — a custom loss, an unusual data pipeline, a multi-stage job that is only partly training — Modal will run it and a template-driven platform will not. Their per-second billing is also genuinely excellent.

Model deployment and serving with strong operational tooling, compliance posture and cold-start engineering.

Pricing. Published per minute: H100 80GB at $0.10833/min, which is $6.50 per GPU-hour; A100 80GB at $0.06667/min, or $4.00 per hour. Training is offered but is not priced per training token. SOC 2 Type II and HIPAA are on the base plan.

Strongest at. Production operations. Cold-start work, compliance certifications on the entry plan, and support that engages at an engineering level rather than a ticket level. If you are deploying into a regulated environment, that is worth more than a lower hourly rate.

A catalogue of community-published models behind one API, billed per second of compute, with fine-tuning on a subset.

Pricing. Per-second GPU billing: H100 at $0.001525/sec, which is $5.49 per GPU-hour; A100 80GB at $5.04 and L40S at $3.51 per hour. Fine-tuning is billed as compute time on the same rates rather than per training token.

Strongest at. Breadth beyond language models and the shortest possible path from idea to running inference. If you want image, video or audio models alongside text, Replicate covers ground nobody in this set matches.

Managed Ray. Distributed training and serving for teams that have outgrown a single machine.

Pricing. Billed as a platform fee on top of underlying cloud instance cost, which varies by the account it runs in; no standalone GPU-hour rate is published.

Strongest at. Genuine distributed workloads. If a run does not fit on one node and the parallelism is not embarrassingly simple, Ray is the right abstraction and Anyscale is the managed version of it.

Closed model API

Fine-tuning of OpenAI’s own closed models, served only from OpenAI. The default first stop, and the thing most teams are trying to leave.

Pricing. Priced per training token and per inference token by model, with fine-tuned inference charged above the base rate. Rates are not reproduced here because the pricing page could not be retrieved at the date checked; see the source for current figures.

Strongest at. Quality per unit of effort. The base models are strong, the fine-tuning API is two calls, and nobody has to think about VRAM. For a team without ML experience and without a portability requirement, that is a real advantage.

A managed API over several model vendors, with custom-model fine-tuning and provisioned throughput inside AWS.

Pricing. Custom models are billed per training token plus provisioned model units per hour for serving; no per-GPU-hour rate exists to compare.

Strongest at. One API across several model vendors, inside an AWS account, with the governance story already written. For an enterprise that needs to switch model providers without switching procurement, that is the whole value proposition.

Hyperscaler ML platform

The full machine-learning platform on AWS. Enormously capable, correspondingly heavy, and priced per instance-hour.

Pricing. Billed per instance-hour by instance family and region, with H100 capacity sold as multi-GPU instances rather than single cards, so no per-GPU-hour figure is directly comparable.

Strongest at. It is already approved. If your organisation has an AWS agreement, a security review that took nine months and data that is not allowed to leave the account, none of the rest of this comparison matters.

Google Cloud’s machine-learning platform, with tuning for Gemini and for a set of open-weight models.

Pricing. Priced per node-hour and per token depending on the product path, varying by region; there is no single comparable H100 hourly rate.

Strongest at. If your data is already in BigQuery, the distance from data to trained model is shorter here than anywhere else, and that distance is usually where projects actually die.

Last verified 6 August 2026. Every figure was read from the vendor's own pricing page on the date shown, and each row links to it.

Start with the free tier

A magic link creates your account, your tenant and your first API key. No card until you ask for compute.