Capabilities and Hard Limits
What LLMs do well, what they cannot do by construction, and why the frontier is jagged — testing beats trusting.
Content last verified 2026-09.
Lessons
Sources
- Brown et al. (2020) — Language Models are Few-Shot Learners (GPT-3)
- Wei et al. (2022) — Emergent Abilities of Large Language Models
- Schaeffer, Miranda & Koyejo (2023) — Are Emergent Abilities of Large Language Models a Mirage?
- Dell’Acqua et al. (2023) — Navigating the Jagged Technological Frontier (SSRN working paper)
- Liang et al. (2022) — Holistic Evaluation of Language Models (HELM)