Next-Token Prediction, Really
The single objective behind everything an LLM does: a probability for every token, a measurable notion of being wrong, and how prediction becomes ability.
Content last verified 2026-09.
Lessons
- The Guess-the-Next-Token Game
- A Probability for Every Token
- Being Wrong, Measurably: Loss
- How Prediction Becomes Ability
Sources
- Shannon (1948) — A Mathematical Theory of Communication (entropy)
- Shannon (1951) — Prediction and Entropy of Printed English
- Brown et al. (2020) — Language Models are Few-Shot Learners (GPT-3, in-context learning)
- Wei et al. (2022) — Emergent Abilities of Large Language Models
- Schaeffer, Miranda & Koyejo (2023) — Are Emergent Abilities of Large Language Models a Mirage?