Data & representations · Glossary term
What is Data Deduplication?
Detecting and removing exact and near-duplicate examples within or across datasets.
Why does Data Deduplication matter?
Repetition can distort the training distribution, increase memorization, leak test material, and make evaluation appear stronger than it is.
Data Deduplication in practice
Normalize content, use exact hashes and similarity methods, review borderline clusters, and record which version and rule removed each example.
What is the common confusion about Data Deduplication?
Deduplication is not ordinary data cleaning. Two distinct records can legitimately share text, and two paraphrases can still carry the same leaked information.
Learn Data Deduplication in the course
No lesson links to this term yet. Search the course catalog for it.
Related terms
- Data ProvenanceTraceable information about where data originated, who or what transformed it, which versions were used, and how derived artifacts relate…
- Benchmark ContaminationOverlap or information leakage between evaluation examples and data used to pretrain, tune, prompt, select, or otherwise improve the…
- Dataset SplitA documented partition of examples into separate subsets for fitting, development decisions, and final evaluation.
- OverfittingA generalization gap in which performance on training data is substantially better than performance on representative unseen data.
Sources
More terms in Data & representations
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.