Data Pipelines and Curation

Where training text actually comes from, and the pipeline that turns a crawl into a corpus: filtering, deduplication, and the mixture decisions that shape a model.

Content last verified 2026-09.

Lessons

  1. Where the Data Comes From
  2. Cleaning and Filtering
  3. Deduplication: The Unglamorous Win
  4. The Mixture: What the Model Eats

Sources