In Production: Should You Ever Pre-train?
The domain’s production capstone: the honest decision framework for training from scratch versus continued pre-training versus not doing this at all — and what a training cluster demands.
Content last verified 2026-09.
Lessons
- Almost Never — and the Exceptions
- What a Training Cluster Demands
- Training Infrastructure on the Clouds
Sources
- Zhang et al. (2022) — OPT: Open Pre-trained Transformer Language Models (with public training logbook)
- Hoffmann et al. (2022) — Training Compute-Optimal Large Language Models (Chinchilla)
- Touvron et al. (2023) — LLaMA: Open and Efficient Foundation Language Models
- Shoeybi et al. (2019) — Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism