onrup

Serverless GPU platform

Onrup vs Replicate

Should I use Onrup or Replicate?

Replicate is the better choice if you need image, video or audio models alongside text, or want to run somebody else’s published model in one call. Onrup is the better choice if you are fine-tuning language models on your own data and want the run to be predictable and gated.

A catalogue of community-published models behind one API, billed per second of compute, with fine-tuning on a subset.

Where they differ

OnrupReplicate
ModalitiesLanguage models onlyText, image, video and audio
Model sourceA curated catalogue of 43 base modelsA community catalogue of published models
H100 rate$2.99 per GPU-hour serving$5.49 per GPU-hour
Fine-tuningThe core product, with an evaluation gateAvailable on a subset of models
Weight portabilityAlways downloadableModel-dependent

On comparable hardware

Replicate publishes $5.49 per H100 80GB hour. Our serving rate for the same class is $2.99, which makes theirs 1.8× the rate. Total cost is the rate multiplied by wall time, so this ratio is the starting point of a comparison rather than the end of one.

Per-second GPU billing: H100 at $0.001525/sec, which is $5.49 per GPU-hour; A100 80GB at $5.04 and L40S at $3.51 per hour. Fine-tuning is billed as compute time on the same rates rather than per training token. Source, checked 2026-08-06.

The longer answer

Replicate’s catalogue is its argument. Thousands of published models across every modality, each runnable in a single API call, is a genuinely different proposition from a curated list of base models to fine-tune. If your product needs to generate an image and summarise a document, one of those catalogues covers you and the other does not.

The trade is depth for breadth. Fine-tuning on Replicate is available on a subset of models and is billed as compute time rather than being the organising idea of the product. There is no evaluation gate, no cost authorisation before a run, and weight portability varies by model.

For language-model fine-tuning specifically, the comparison is straightforward on price — $5.49 against $2.99 per H100-hour — and on process. For anything beyond language, we are not a candidate at all.

Where Replicate wins

Breadth beyond language models and the shortest possible path from idea to running inference. If you want image, video or audio models alongside text, Replicate covers ground nobody in this set matches.

Choose them if

  • You need image, video or audio models as well as text
  • You want to try somebody else’s published model in one API call
  • Your usage is occasional and per-second billing suits it

On ownership

Weights on Onrup are downloadable from every finished run and publishable to a model hub in one call. On Replicate: Fine-tune outputs are downloadable for many, not all, models.

Frequently asked questions

Do you support image or video models?

No. The catalogue is language models only, and there is no plan to change that. Replicate is the right answer for other modalities.

Can I publish a model for others to run?

You can publish to a model hub, which makes it downloadable and usable by anyone. What we do not offer is a hosted marketplace where others call your model and you earn from it.

Is per-second billing available?

Yes — compute is metered per GPU-second above a one-minute minimum, so a short job is billed as a short job.

Researching Replicate alternatives more broadly? →

Last verified 6 August 2026. Replicate figures come from their own pricing page on the date checked.

Start with the free tier

A magic link creates your account, your tenant and your first API key. No card until you ask for compute.