Data
Deduplication
What is deduplication?
Deduplication removes repeated or near-repeated examples from a training set. Duplicates effectively increase the weight of the repeated content, causing memorisation and distorting evaluation when the same item appears in both training and test splits.
Exact duplicates are easy to find and are usually an artefact of assembly — the same record exported twice, or a join that fanned out. Near-duplicates are harder and more common in real data.
The most damaging case is a duplicate that spans the training and held-out splits, because the evaluation then measures memorisation and reports it as quality. Deduplicate before splitting, not after.
Related terms
Start with the free tier
A magic link creates your account, your tenant and your first API key. No card until you ask for compute.