# Onrup > Fine-tune open-weight models, prove they beat the baseline, and serve them on an OpenAI-compatible endpoint. Serverless compute across multiple providers. Onrup is an API-first platform for fine-tuning open-weight language models and serving them on an OpenAI-compatible endpoint. Compute is serverless across multiple providers; a caller chooses a GPU class and a scaling mode rather than a machine. Three properties define the product: cost is authorised and reserved before compute is leased, a blinded evaluation gate must pass before a model can be deployed or published, and trained weights are always downloadable. Deliberately not published: which compute providers are used, how capacity is selected between them, and how the platform is deployed. Everything a caller can observe — GPU class, memory, hourly rate, scaling behaviour, what is metered and when — is published in full and linked below. Contact: hello@onrup.com ## Start here - [How it works](https://www.onrup.com/how-it-works): dataset to deployed model in five API calls. - [Pricing](https://www.onrup.com/pricing): published rate card, 13 GPU classes, from $0.09 per GPU-hour. - [Quickstart](https://www.onrup.com/docs/quickstart): every call, in full. - [FAQ](https://www.onrup.com/faq): including the questions where the answer is no. ## What makes it different - [Predictable cost](https://www.onrup.com/predictable-cost): spend is reserved against a limit before compute is leased, not alerted on afterwards. - [Your weights, your endpoint](https://www.onrup.com/your-weights): download or publish any finished run; serving speaks the OpenAI API. - [The evaluation gate](https://www.onrup.com/evaluation-gate): blinded comparison against the incumbent; pass, fail or inconclusive. - [Serverless compute](https://www.onrup.com/serverless-compute): a GPU class, not a machine, across multiple providers. ## Documentation - [Quickstart](https://www.onrup.com/docs/quickstart): How do I fine-tune and deploy a model on Onrup? - [Authentication](https://www.onrup.com/docs/authentication): How does authentication work on the Onrup API? - [Datasets](https://www.onrup.com/docs/datasets): What dataset formats does Onrup accept and how are they validated? - [Runs](https://www.onrup.com/docs/runs): How does a training run work on Onrup? - [Evaluation and the gate](https://www.onrup.com/docs/evaluation): How does the Onrup evaluation gate decide whether a model can be deployed? - [Deployments](https://www.onrup.com/docs/deployments): How do I serve a fine-tuned model on Onrup? - [Publishing models](https://www.onrup.com/docs/publishing): How do I export or publish a model trained on Onrup? - [Long-running operations](https://www.onrup.com/docs/operations): How do I track a long-running job on the Onrup API? - [Billing and spend limits](https://www.onrup.com/docs/billing): How does Onrup billing work? - [Errors](https://www.onrup.com/docs/errors): What error codes does the Onrup API return? - [Rate limits and quotas](https://www.onrup.com/docs/limits): What are the Onrup rate limits and quotas? - [Security model](https://www.onrup.com/docs/security): How does Onrup isolate and protect tenant data? - [API reference](https://www.onrup.com/docs/api): What endpoints does the Onrup API expose? ## Model catalogue (43 models, 12 families) - [SmolLM2 1.7B](https://www.onrup.com/models/smollm2-1.7b): 1.7B, 6 GB LoRA / 4 GB QLoRA, Apache 2.0. - [SmolLM3 3B](https://www.onrup.com/models/smollm3-3b): 3B, 10 GB LoRA / 6 GB QLoRA, Apache 2.0. - [Qwen3 0.6B](https://www.onrup.com/models/qwen3-0.6b): 0.6B, 4 GB LoRA / 3 GB QLoRA, Apache 2.0. - [Qwen3 1.7B](https://www.onrup.com/models/qwen3-1.7b): 1.7B, 6 GB LoRA / 4 GB QLoRA, Apache 2.0. - [Qwen3 4B](https://www.onrup.com/models/qwen3-4b): 4B, 12 GB LoRA / 8 GB QLoRA, Apache 2.0. - [Qwen3 8B](https://www.onrup.com/models/qwen3-8b): 8B, 18 GB LoRA / 14 GB QLoRA, Apache 2.0. - [Qwen3 14B](https://www.onrup.com/models/qwen3-14b): 14B, 28 GB LoRA / 20 GB QLoRA, Apache 2.0. - [Qwen3 32B](https://www.onrup.com/models/qwen3-32b): 32B, 64 GB LoRA / 48 GB QLoRA, Apache 2.0. - [Qwen3 30B-A3B](https://www.onrup.com/models/qwen3-30b-a3b): 30B, 64 GB LoRA / 36 GB QLoRA, Apache 2.0. - [Llama 3.2 1B](https://www.onrup.com/models/llama-3.2-1b): 1B, 4 GB LoRA / 3 GB QLoRA, Llama 3.2 Community. - [Llama 3.2 3B](https://www.onrup.com/models/llama-3.2-3b): 3B, 10 GB LoRA / 6 GB QLoRA, Llama 3.2 Community. - [Llama 3.1 8B](https://www.onrup.com/models/llama-3.1-8b): 8B, 18 GB LoRA / 14 GB QLoRA, Llama 3.1 Community. - [Llama 3.3 70B](https://www.onrup.com/models/llama-3.3-70b): 70B, 140 GB LoRA / 48 GB QLoRA, Llama 3.3 Community. - [Llama 4 Scout 109B](https://www.onrup.com/models/llama-4-scout-109b): 109B, 220 GB LoRA / 80 GB QLoRA, Llama 4 Community. - [Mistral 7B v0.3](https://www.onrup.com/models/mistral-7b-v0.3): 7B, 16 GB LoRA / 12 GB QLoRA, Apache 2.0. - [Mistral Nemo 12B](https://www.onrup.com/models/mistral-nemo-12b): 12B, 24 GB LoRA / 16 GB QLoRA, Apache 2.0. - [Mistral Small 3 24B](https://www.onrup.com/models/mistral-small-3-24b): 24B, 48 GB LoRA / 36 GB QLoRA, Apache 2.0. - [Phi-4 Mini 3.8B](https://www.onrup.com/models/phi-4-mini): 3.8B, 12 GB LoRA / 8 GB QLoRA, MIT. - [Phi-4 14B](https://www.onrup.com/models/phi-4-14b): 14B, 28 GB LoRA / 20 GB QLoRA, MIT. - [Gemma 3 1B](https://www.onrup.com/models/gemma-3-1b): 1B, 4 GB LoRA / 3 GB QLoRA, Gemma Terms of Use. - [Gemma 3 4B](https://www.onrup.com/models/gemma-3-4b): 4B, 12 GB LoRA / 8 GB QLoRA, Gemma Terms of Use. - [Gemma 3 12B](https://www.onrup.com/models/gemma-3-12b): 12B, 24 GB LoRA / 18 GB QLoRA, Gemma Terms of Use. - [Gemma 3 27B](https://www.onrup.com/models/gemma-3-27b): 27B, 56 GB LoRA / 36 GB QLoRA, Gemma Terms of Use. - [DeepSeek-R1 Distill Qwen 7B](https://www.onrup.com/models/deepseek-r1-distill-qwen-7b): 7B, 16 GB LoRA / 12 GB QLoRA, Apache 2.0. - [DeepSeek-R1 Distill Llama 8B](https://www.onrup.com/models/deepseek-r1-distill-llama-8b): 8B, 18 GB LoRA / 14 GB QLoRA, Llama 3.1 Community. - [LFM2 350M](https://www.onrup.com/models/lfm2-350m): 0.35B, 3 GB LoRA / 2 GB QLoRA, LFM Open. - [LFM2 700M](https://www.onrup.com/models/lfm2-700m): 0.7B, 4 GB LoRA / 3 GB QLoRA, LFM Open. - [LFM2 1.2B](https://www.onrup.com/models/lfm2-1.2b): 1.2B, 5 GB LoRA / 4 GB QLoRA, LFM Open. - [LFM2 2.6B](https://www.onrup.com/models/lfm2-2.6b): 2.6B, 8 GB LoRA / 6 GB QLoRA, LFM Open. - [LFM2 8B-A1B](https://www.onrup.com/models/lfm2-8b-moe): 8.3B, 18 GB LoRA / 12 GB QLoRA, LFM Open. - [LFM2 24B-A2B](https://www.onrup.com/models/lfm2-24b-a2b): 24B, 48 GB LoRA / 28 GB QLoRA, LFM Open. - [Granite 3.0 2B](https://www.onrup.com/models/granite-3.0-2b): 2B, 7 GB LoRA / 5 GB QLoRA, Apache 2.0. - [Granite 3.0 8B](https://www.onrup.com/models/granite-3.0-8b): 8B, 18 GB LoRA / 14 GB QLoRA, Apache 2.0. - [Granite 4.1 8B](https://www.onrup.com/models/granite-4.1-8b): 8B, 18 GB LoRA / 14 GB QLoRA, Apache 2.0. - [gpt-oss 20B](https://www.onrup.com/models/gpt-oss-20b): 20.9B, 40 GB LoRA / 24 GB QLoRA, Apache 2.0. - [gpt-oss 120B](https://www.onrup.com/models/gpt-oss-120b): 116.8B, 220 GB LoRA / 80 GB QLoRA, Apache 2.0. - [Falcon 3 1B](https://www.onrup.com/models/falcon-3-1b): 1B, 4 GB LoRA / 3 GB QLoRA, TII Falcon LLM. - [Falcon 3 3B](https://www.onrup.com/models/falcon-3-3b): 3B, 10 GB LoRA / 6 GB QLoRA, TII Falcon LLM. - [Falcon 3 7B](https://www.onrup.com/models/falcon-3-7b): 7B, 16 GB LoRA / 12 GB QLoRA, TII Falcon LLM. - [Falcon 3 10B](https://www.onrup.com/models/falcon-3-10b): 10B, 22 GB LoRA / 16 GB QLoRA, TII Falcon LLM. - [OLMo 2 7B](https://www.onrup.com/models/olmo-2-7b): 7B, 16 GB LoRA / 12 GB QLoRA, Apache 2.0. - [OLMo 3 7B](https://www.onrup.com/models/olmo-3-7b): 7B, 16 GB LoRA / 12 GB QLoRA, Apache 2.0. - [OLMo 3 32B](https://www.onrup.com/models/olmo-3-32b): 32B, 64 GB LoRA / 48 GB QLoRA, Apache 2.0. ## GPU classes (13) - [RTX 3080](https://www.onrup.com/gpus/rtx-3080-12g): 12 GB, $0.09/hr training, $0.11/hr serving. - [RTX 4000 Ada](https://www.onrup.com/gpus/rtx-4000-ada-20g): 20 GB, $0.09/hr training, $0.11/hr serving. - [L4](https://www.onrup.com/gpus/l4-24g): 24 GB, $0.17/hr training, $0.20/hr serving. - [RTX 3090](https://www.onrup.com/gpus/rtx-3090-24g): 24 GB, $0.22/hr training, $0.26/hr serving. - [RTX 4090](https://www.onrup.com/gpus/rtx-4090-24g): 24 GB, $0.38/hr training, $0.43/hr serving. - [A40](https://www.onrup.com/gpus/a40-48g): 48 GB, $0.42/hr training, $0.48/hr serving. - [RTX 6000 Ada](https://www.onrup.com/gpus/rtx-6000-ada-48g): 48 GB, $0.61/hr training, $0.71/hr serving. - [L40S](https://www.onrup.com/gpus/l40s-48g): 48 GB, $0.78/hr training, $0.90/hr serving. - [A6000](https://www.onrup.com/gpus/a6000-48g): 48 GB, $0.65/hr training, $0.75/hr serving. - [A100 40 GB](https://www.onrup.com/gpus/a100-40g): 40 GB, $1.17/hr training, $1.35/hr serving. - [A100 80 GB](https://www.onrup.com/gpus/a100-80g): 80 GB, $1.68/hr training, $1.94/hr serving. - [H100 80 GB](https://www.onrup.com/gpus/h100-80g): 80 GB, $2.59/hr training, $2.99/hr serving. - [H200](https://www.onrup.com/gpus/h200-141g): 141 GB, $4.55/hr training, $5.25/hr serving. ## Guides - [Should you fine-tune, or is it a prompt problem?](https://www.onrup.com/guides/should-you-fine-tune): When is fine-tuning worth it compared to prompt engineering or retrieval? - [LoRA, QLoRA or full fine-tuning: how to choose](https://www.onrup.com/guides/lora-vs-qlora-vs-full-fine-tuning): What is the difference between LoRA, QLoRA and full fine-tuning? - [How much VRAM does fine-tuning actually need?](https://www.onrup.com/guides/how-much-vram-do-i-need): How much GPU memory do I need to fine-tune a model? - [Preparing a fine-tuning dataset](https://www.onrup.com/guides/preparing-a-fine-tuning-dataset): How do I prepare a dataset for fine-tuning? - [ShareGPT, ChatML and Alpaca: which dataset format to use](https://www.onrup.com/guides/sharegpt-vs-chatml-vs-alpaca): What is the difference between ShareGPT, ChatML and Alpaca dataset formats? - [How to choose a base model](https://www.onrup.com/guides/choosing-a-base-model): How do I choose which base model to fine-tune? - [Designing an evaluation set that actually decides things](https://www.onrup.com/guides/designing-an-evaluation-set): How do I build an evaluation set for a fine-tuned model? - [How to estimate what a fine-tuning run will cost](https://www.onrup.com/guides/estimating-fine-tuning-cost): How much does it cost to fine-tune an open-weight model? - [Scale to zero or always warm: choosing a serving mode](https://www.onrup.com/guides/scale-to-zero-or-always-warm): Should my model endpoint scale to zero or stay always warm? - [Migrating from a closed model API to an open-weight fine-tune](https://www.onrup.com/guides/migrating-from-a-closed-model-api): How do I replace a closed model API with a fine-tuned open-weight model? - [Why fine-tuned models fail, and how to tell which failure you have](https://www.onrup.com/guides/why-fine-tuned-models-fail): Why did my fine-tuned model get worse instead of better? - [When to use preference tuning instead of supervised fine-tuning](https://www.onrup.com/guides/when-to-use-preference-tuning): When should I use DPO instead of supervised fine-tuning? - [Using reinforcement learning to improve reasoning](https://www.onrup.com/guides/reinforcement-learning-for-reasoning): How does GRPO work and when should I use reinforcement learning to fine-tune? - [Serving many fine-tuned models without paying for each one](https://www.onrup.com/guides/serving-many-fine-tuned-models): How can I serve dozens of fine-tuned models economically? - [Controlling spend on a fine-tuning platform](https://www.onrup.com/guides/controlling-fine-tuning-spend): How do I stop fine-tuning and inference costs running away? - [Keeping a fine-tuned model portable](https://www.onrup.com/guides/keeping-a-fine-tuned-model-portable): How do I make sure I can move a fine-tuned model to another provider? ## Use cases - [Customer support triage](https://www.onrup.com/use-cases/customer-support-triage): How do I fine-tune a model to route and tag support tickets? - [Structured data extraction](https://www.onrup.com/use-cases/structured-data-extraction): How do I fine-tune a model to return reliable JSON from messy documents? - [Replacing a frontier API on one task](https://www.onrup.com/use-cases/replacing-a-frontier-api): How do I cut my LLM bill by replacing a frontier model with a fine-tuned open one? - [Domain question answering](https://www.onrup.com/use-cases/domain-question-answering): Should I fine-tune or use retrieval for questions about my own documents? - [Code review assistant](https://www.onrup.com/use-cases/code-review-assistant): How do I fine-tune a model on my team’s code review standards? - [SQL generation over your own schema](https://www.onrup.com/use-cases/sql-generation): How do I fine-tune a model to write correct SQL for my database? - [Tool-calling agents](https://www.onrup.com/use-cases/tool-calling-agent): How do I fine-tune a model to call my tools reliably? - [Content moderation](https://www.onrup.com/use-cases/content-moderation): How do I fine-tune a moderation model on my own policy? - [Summarisation in a house style](https://www.onrup.com/use-cases/summarisation-with-house-style): How do I fine-tune a model to summarise in our specific format? - [PII redaction](https://www.onrup.com/use-cases/pii-redaction): How do I fine-tune a model to redact personal data in our documents? - [Reasoning with reinforcement learning](https://www.onrup.com/use-cases/reasoning-with-rl): When should I use GRPO instead of supervised fine-tuning? - [Translation and localisation](https://www.onrup.com/use-cases/translation-and-localisation): How do I fine-tune a model on our terminology and tone across languages? - [Classification at volume](https://www.onrup.com/use-cases/classification-at-volume): What is the cheapest way to classify millions of documents with an LLM? - [On-brand copy generation](https://www.onrup.com/use-cases/on-brand-copy-generation): How do I fine-tune a model to write in our brand voice? - [Semantic routing between models](https://www.onrup.com/use-cases/semantic-routing): How do I route requests between a small and a large model automatically? - [Document parsing pipelines](https://www.onrup.com/use-cases/document-parsing-pipelines): How do I fine-tune a model to parse messy PDFs and scans consistently? ## Comparisons - [The landscape](https://www.onrup.com/landscape): sourced market map; every figure links to the vendor's own page. - [Onrup vs Together AI](https://www.onrup.com/compare/onrup-vs-together-ai): A broad model API with fine-tuning attached, covering the widest catalogue of open-weight models of anyone in this set. - [Onrup vs Fireworks AI](https://www.onrup.com/compare/onrup-vs-fireworks-ai): Inference, fine-tuning and evaluation behind one API, with fine-tuned adapters served at the same per-token rate as the base model. - [Onrup vs Modal](https://www.onrup.com/compare/onrup-vs-modal): Serverless compute for arbitrary Python, billed by the second. Not a fine-tuning product — a substrate you build one on. - [Onrup vs Baseten](https://www.onrup.com/compare/onrup-vs-baseten): Model deployment and serving with strong operational tooling, compliance posture and cold-start engineering. - [Onrup vs Replicate](https://www.onrup.com/compare/onrup-vs-replicate): A catalogue of community-published models behind one API, billed per second of compute, with fine-tuning on a subset. - [Onrup vs OpenAI fine-tuning](https://www.onrup.com/compare/onrup-vs-openai-fine-tuning): Fine-tuning of OpenAI’s own closed models, served only from OpenAI. The default first stop, and the thing most teams are trying to leave. - [Onrup vs Hugging Face](https://www.onrup.com/compare/onrup-vs-hugging-face): The model hub itself, plus AutoTrain for training and Inference Endpoints for serving. The centre of gravity of the open-weight world. - [Onrup vs Predibase](https://www.onrup.com/compare/onrup-vs-predibase): A fine-tuning platform built around serving many LoRA adapters from a single GPU. Acquired — predibase.com now redirects to Rubrik. - [Onrup vs OpenPipe](https://www.onrup.com/compare/onrup-vs-openpipe): Capture production traffic from a large model, then train a small one to replace it on that exact distribution. - [Onrup vs AWS SageMaker](https://www.onrup.com/compare/onrup-vs-aws-sagemaker): The full machine-learning platform on AWS. Enormously capable, correspondingly heavy, and priced per instance-hour. - [Onrup vs Google Vertex AI](https://www.onrup.com/compare/onrup-vs-google-vertex-ai): Google Cloud’s machine-learning platform, with tuning for Gemini and for a set of open-weight models. - [Onrup vs AWS Bedrock](https://www.onrup.com/compare/onrup-vs-aws-bedrock): A managed API over several model vendors, with custom-model fine-tuning and provisioned throughput inside AWS. - [Onrup vs Anyscale](https://www.onrup.com/compare/onrup-vs-anyscale): Managed Ray. Distributed training and serving for teams that have outgrown a single machine. - [Onrup vs Lamini](https://www.onrup.com/compare/onrup-vs-lamini): Enterprise fine-tuning with a focus on factual accuracy and hallucination reduction, deployable on-premises. ## Integrations - [OpenAI Python SDK](https://www.onrup.com/integrations/openai-python): Python. - [OpenAI Node SDK](https://www.onrup.com/integrations/openai-node): TypeScript. - [curl and plain HTTP](https://www.onrup.com/integrations/curl): Shell. - [LangChain](https://www.onrup.com/integrations/langchain): Python. - [LlamaIndex](https://www.onrup.com/integrations/llamaindex): Python. - [Vercel AI SDK](https://www.onrup.com/integrations/vercel-ai-sdk): TypeScript. - [LiteLLM](https://www.onrup.com/integrations/litellm): YAML. - [Pydantic AI](https://www.onrup.com/integrations/pydantic-ai): Python. - [n8n](https://www.onrup.com/integrations/n8n): Configuration. - [Open WebUI](https://www.onrup.com/integrations/open-webui): Configuration. - [Hugging Face Hub](https://www.onrup.com/integrations/hugging-face-hub): HTTP. - [Continue](https://www.onrup.com/integrations/continue): JSON. ## Glossary (124 terms) - [Fine-tuning](https://www.onrup.com/glossary/fine-tuning) - [LoRA](https://www.onrup.com/glossary/lora) - [QLoRA](https://www.onrup.com/glossary/qlora) - [Supervised fine-tuning](https://www.onrup.com/glossary/supervised-fine-tuning) - [Preference tuning](https://www.onrup.com/glossary/preference-tuning) - [Reinforcement learning from verifiable rewards](https://www.onrup.com/glossary/reinforcement-learning-from-verifiable-rewards) - [Full fine-tuning](https://www.onrup.com/glossary/full-fine-tuning) - [Catastrophic forgetting](https://www.onrup.com/glossary/catastrophic-forgetting) - [Overfitting](https://www.onrup.com/glossary/overfitting) - [Epoch](https://www.onrup.com/glossary/epoch) - [Learning rate](https://www.onrup.com/glossary/learning-rate) - [Learning rate schedule](https://www.onrup.com/glossary/learning-rate-schedule) - [Warmup](https://www.onrup.com/glossary/warmup) - [Batch size](https://www.onrup.com/glossary/batch-size) - [Gradient accumulation](https://www.onrup.com/glossary/gradient-accumulation) - [Gradient checkpointing](https://www.onrup.com/glossary/gradient-checkpointing) - [Rank](https://www.onrup.com/glossary/rank) - [Target modules](https://www.onrup.com/glossary/target-modules) - [Adapter](https://www.onrup.com/glossary/adapter) - [Adapter merging](https://www.onrup.com/glossary/adapter-merging) - [Training step](https://www.onrup.com/glossary/training-step) - [Checkpoint](https://www.onrup.com/glossary/checkpoint) - [Early stopping](https://www.onrup.com/glossary/early-stopping) - [Optimiser state](https://www.onrup.com/glossary/optimiser-state) - [Gradient](https://www.onrup.com/glossary/gradient) - [Loss](https://www.onrup.com/glossary/loss) - [Perplexity](https://www.onrup.com/glossary/perplexity) - [Training example](https://www.onrup.com/glossary/training-example) - [ChatML](https://www.onrup.com/glossary/chatml) - [ShareGPT format](https://www.onrup.com/glossary/sharegpt-format) - [Alpaca format](https://www.onrup.com/glossary/alpaca-format) - [Format adapter](https://www.onrup.com/glossary/format-adapter) - [Dataset validation](https://www.onrup.com/glossary/dataset-validation) - [Deduplication](https://www.onrup.com/glossary/deduplication) - [Held-out set](https://www.onrup.com/glossary/held-out-set) - [Validation split](https://www.onrup.com/glossary/validation-split) - [Data contamination](https://www.onrup.com/glossary/data-contamination) - [Synthetic data](https://www.onrup.com/glossary/synthetic-data) - [Data provenance](https://www.onrup.com/glossary/data-provenance) - [Sequence length](https://www.onrup.com/glossary/sequence-length) - [Truncation](https://www.onrup.com/glossary/truncation) - [Tokenisation](https://www.onrup.com/glossary/tokenisation) - [Token](https://www.onrup.com/glossary/token) - [Special tokens](https://www.onrup.com/glossary/special-tokens) - [Chat template](https://www.onrup.com/glossary/chat-template) - [Preference pair](https://www.onrup.com/glossary/preference-pair) - [Evaluation gate](https://www.onrup.com/glossary/evaluation-gate) - [Blinded evaluation](https://www.onrup.com/glossary/blinded-evaluation) - [Baseline model](https://www.onrup.com/glossary/baseline-model) - [Win rate](https://www.onrup.com/glossary/win-rate) - [Confidence interval](https://www.onrup.com/glossary/confidence-interval) - [Statistical significance](https://www.onrup.com/glossary/statistical-significance) - [LLM-as-judge](https://www.onrup.com/glossary/llm-as-judge) - [Rubric](https://www.onrup.com/glossary/rubric) - [Critical regression](https://www.onrup.com/glossary/critical-regression) - [Regression](https://www.onrup.com/glossary/regression) - [Benchmark](https://www.onrup.com/glossary/benchmark) - [Safety evaluation](https://www.onrup.com/glossary/safety-evaluation) - [Inference](https://www.onrup.com/glossary/inference) - [OpenAI-compatible API](https://www.onrup.com/glossary/openai-compatible-api) - [Endpoint](https://www.onrup.com/glossary/endpoint) - [Scale to zero](https://www.onrup.com/glossary/scale-to-zero) - [Always warm](https://www.onrup.com/glossary/always-warm) - [Cold start](https://www.onrup.com/glossary/cold-start) - [Cooldown](https://www.onrup.com/glossary/cooldown) - [Replica](https://www.onrup.com/glossary/replica) - [Latency](https://www.onrup.com/glossary/latency) - [Time to first token](https://www.onrup.com/glossary/time-to-first-token) - [Throughput](https://www.onrup.com/glossary/throughput) - [Batching](https://www.onrup.com/glossary/batching) - [KV cache](https://www.onrup.com/glossary/kv-cache) - [Context window](https://www.onrup.com/glossary/context-window) - [Streaming](https://www.onrup.com/glossary/streaming) - [Server-sent events](https://www.onrup.com/glossary/server-sent-events) - [Multi-adapter serving](https://www.onrup.com/glossary/multi-adapter-serving) - [Quantisation](https://www.onrup.com/glossary/quantisation) - [Tool calling](https://www.onrup.com/glossary/tool-calling) - [Structured output](https://www.onrup.com/glossary/structured-output) - [Schema validity](https://www.onrup.com/glossary/schema-validity) - [Base model](https://www.onrup.com/glossary/base-model) - [Instruction-tuned model](https://www.onrup.com/glossary/instruction-tuned-model) - [Open-weight model](https://www.onrup.com/glossary/open-weight-model) - [Model licence](https://www.onrup.com/glossary/model-licence) - [Mixture of experts](https://www.onrup.com/glossary/mixture-of-experts) - [Active parameters](https://www.onrup.com/glossary/active-parameters) - [Dense model](https://www.onrup.com/glossary/dense-model) - [Parameter count](https://www.onrup.com/glossary/parameter-count) - [Attention](https://www.onrup.com/glossary/attention) - [Transformer](https://www.onrup.com/glossary/transformer) - [Feed-forward network](https://www.onrup.com/glossary/feed-forward-network) - [Reasoning model](https://www.onrup.com/glossary/reasoning-model) - [Distillation](https://www.onrup.com/glossary/distillation) - [Retrieval-augmented generation](https://www.onrup.com/glossary/retrieval-augmented-generation) - [Hallucination](https://www.onrup.com/glossary/hallucination) - [Refusal](https://www.onrup.com/glossary/refusal) - [Prompt injection](https://www.onrup.com/glossary/prompt-injection) - [Temperature](https://www.onrup.com/glossary/temperature) - [Agent](https://www.onrup.com/glossary/agent) - [GPU class](https://www.onrup.com/glossary/gpu-class) - [VRAM](https://www.onrup.com/glossary/vram) - [GPU-hour](https://www.onrup.com/glossary/gpu-hour) - [Serverless compute](https://www.onrup.com/glossary/serverless-compute) - [Cost authorisation](https://www.onrup.com/glossary/cost-authorisation) - [Spend limit](https://www.onrup.com/glossary/spend-limit) - [Reservation](https://www.onrup.com/glossary/reservation) - [Metered billing](https://www.onrup.com/glossary/metered-billing) - [Usage event](https://www.onrup.com/glossary/usage-event) - [Entitlement](https://www.onrup.com/glossary/entitlement) - [Quota](https://www.onrup.com/glossary/quota) - [Rate limit](https://www.onrup.com/glossary/rate-limit) - [Durable operation](https://www.onrup.com/glossary/durable-operation) - [Idempotency](https://www.onrup.com/glossary/idempotency) - [Resumption](https://www.onrup.com/glossary/resumption) - [Model weights](https://www.onrup.com/glossary/model-weights) - [Model publishing](https://www.onrup.com/glossary/model-publishing) - [Model card](https://www.onrup.com/glossary/model-card) - [Audit log](https://www.onrup.com/glossary/audit-log) - [API key](https://www.onrup.com/glossary/api-key) - [Scope](https://www.onrup.com/glossary/scope) - [Prompt engineering](https://www.onrup.com/glossary/prompt-engineering) - [Reward function](https://www.onrup.com/glossary/reward-function) - [Reward hacking](https://www.onrup.com/glossary/reward-hacking) - [Reward model](https://www.onrup.com/glossary/reward-model) - [Reference model](https://www.onrup.com/glossary/reference-model) ## Machine-readable - [OpenAPI document](https://www.onrup.com/openapi.json) - [Models as JSON](https://www.onrup.com/api/models.json) - [GPU classes and rates as JSON](https://www.onrup.com/api/pricing.json) - [Glossary as JSON](https://www.onrup.com/api/glossary.json) - [Landscape dataset as JSON](https://www.onrup.com/api/landscape.json) - [Guides as RSS](https://www.onrup.com/rss.xml) - Every guide is also served as Markdown at its URL plus `.md`. ## Legal - [Privacy](https://www.onrup.com/legal/privacy) - [Terms](https://www.onrup.com/legal/terms) - [Acceptable use](https://www.onrup.com/legal/acceptable-use)