Data Contamination
When the test set leaks into the training set: why contamination happens at web scale, how it is detected, and what a score means once it has.
Content last verified 2026-09.
Lessons
Sources
- Brown et al. (2020) — Language Models are Few-Shot Learners (GPT-3; includes the paper’s own benchmark-contamination analysis)
- Dodge et al. (2021) — Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus (C4)
- Lee et al. (2021) — Deduplicating Training Data Makes Language Models Better