Evaluation
Data contamination
What is data contamination?
Data contamination is the presence of evaluation data in the training set, whether directly or through near-duplicates. It inflates measured performance without improving real capability, and it is the most common reason an impressive evaluation fails to survive contact with production.
It happens most often through mundane accidents: deduplication run after splitting rather than before, a shared source document appearing in both halves, or an evaluation set assembled from the same export as the training data.
It also affects public benchmarks, where a model may have seen the test items during pretraining. This is a strong argument for evaluating on your own data rather than on published leaderboards.
Related terms
Start with the free tier
A magic link creates your account, your tenant and your first API key. No card until you ask for compute.