Mixed Precision and Stability
FP32, FP16, BF16 — why training runs in mixed precision, and what loss spikes, checkpoints, and restarts look like at scale.
Content last verified 2026-09.
Lessons
Sources
- Micikevicius et al. (2017) — Mixed Precision Training
- Zhang et al. (2022) — OPT: Open Pre-trained Transformer Language Models (published with a public training logbook)
- Rajbhandari et al. (2019) — ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- Google Cloud TPU documentation — The bfloat16 numerical format
- Chowdhery et al. (2022) — PaLM: Scaling Language Modeling with Pathways (loss-spike mitigation account)