Saturation: When the Ceiling Arrives

Lesson 2 of 4 in Benchmarks and Their Limits.

Benchmarks are born discriminating and die compressed. When a benchmark is new and hard, scores are spread out and a gap between models means something. As models improve, the leaders climb toward the maximum — and the scores cluster. A benchmark where the top contenders all sit within a couple of points of each other, near the top of the scale, has saturated: it can still tell you a model is not terrible, but it can no longer rank the models you are actually choosing between.

This is not a scandal; it is the life cycle of any fixed test. The task set does not get harder, and the models do. The failure mode is not that saturated benchmarks exist — it is that people keep citing them as if the top of a compressed scale still carried ranking information.

Line chart with model generation on the x-axis and benchmark score in percent on the y-axis. Three model-family curves start spread apart between roughly 35 and 50 percent, rise steeply, and converge between 92 and 93 percent by generation six. A fourth flat line labeled as an illustrative effective ceiling from label noise sits at 96 percent, just above where the curves flatten.

How saturation looks: toy scores for three model families across successive generations, compressing toward a ceiling. All numbers are invented for teaching — the shape, not the values, is the point. Once the curves converge near the top, the benchmark stops separating the leaders. (illustrative — source: Saturation of earlier benchmark suites motivated harder ones — Hendrycks et al. (2020))

Near the top, a second effect kicks in: the answer key becomes the ceiling. Large task sets are labeled by humans at scale, and some items end up mislabeled, ambiguous, or genuinely disputed. A model that answered every clean item correctly would still lose points on the flawed ones — so the effective maximum sits below 100, and nobody knows exactly where. When the leaders are within a point or two of each other in that zone, the differences between them can owe as much to which flawed items each happened to match as to any real capability gap.

The community response is a treadmill. When a benchmark saturates, harder successors appear — MMLU itself was proposed after transformer models pushed past earlier natural-language-understanding suites (Hendrycks et al., 2020), and the same fate later drove harder successors to MMLU in turn. The treadmill is healthy — instruments should be replaced when they stop measuring — but it has a side effect you must plan for: scores are not comparable across the treadmill. A model’s number on the old benchmark and a rival’s number on the successor share nothing.

Interactive checkpoint quiz (1 questions) — open this page in a browser to take it.