Data Leakage and Memorization
Models memorize training text and can be made to emit it; prompts flow to providers; deletion is hard. The leakage map, both directions.
Content last verified 2026-09.
Lessons
Sources
- Carlini et al. (2020) — Extracting Training Data from Large Language Models
- Carlini et al. (2022) — Quantifying Memorization Across Neural Language Models
- Lee et al. (2021) — Deduplicating Training Data Makes Language Models Better
- OWASP Top 10 for LLM Applications (2025) — LLM02: Sensitive Information Disclosure