onrup

Evaluation

Data contamination

What is data contamination?

Data contamination is the presence of evaluation data in the training set, whether directly or through near-duplicates. It inflates measured performance without improving real capability, and it is the most common reason an impressive evaluation fails to survive contact with production.

It happens most often through mundane accidents: deduplication run after splitting rather than before, a shared source document appearing in both halves, or an evaluation set assembled from the same export as the training data.

It also affects public benchmarks, where a model may have seen the test items during pretraining. This is a strong argument for evaluating on your own data rather than on published leaderboards.

Related terms

All terms in the glossary →

Start with the free tier

A magic link creates your account, your tenant and your first API key. No card until you ask for compute.