onrup

Evaluation

Benchmark

What is benchmark?

A benchmark is a standardised public evaluation set used to compare models. Benchmarks are useful for shortlisting base models and unreliable for deployment decisions, because they measure general capability rather than performance on your task.

Contamination is endemic. Popular benchmark items appear in the pretraining corpora of the models being tested, which inflates scores in ways that are hard to detect and impossible to correct after the fact.

The useful role is triage: benchmarks narrow a field of forty models to a shortlist of three. The choice among the three should be made on your own data, against your own baseline.

Related terms

All terms in the glossary →

Start with the free tier

A magic link creates your account, your tenant and your first API key. No card until you ask for compute.