Compute Budgets
From FLOPs accounting to GPU-hours: how training budgets are actually estimated, and where the money goes.
Content last verified 2026-09.
Lessons
Sources
- Kaplan et al. (2020) — Scaling Laws for Neural Language Models (C ≈ 6·N·D, §2.1; PF-day unit)
- Hoffmann et al. (2022) — Training Compute-Optimal Large Language Models (Chinchilla)
- Brown et al. (2020) — Language Models are Few-Shot Learners (GPT-3 training scale)
- Zhang et al. (2022) — OPT: Open Pre-trained Transformer Language Models (public training logbook)