Prompt Injection and Jailbreaks: The Model-Level View
Direct and indirect injection, the jailbreak taxonomy, and why refusals are trained dispositions rather than enforcement.
Content last verified 2026-09.
Lessons
Sources
- Anthropic (2024) — Many-shot Jailbreaking
- Greshake et al. (2023) — Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Zou et al. (2023) — Universal and Transferable Adversarial Attacks on Aligned Language Models (GCG)
- Wallace et al. (2024) — The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Ganguli et al. (2022) — Red Teaming Language Models to Reduce Harms
- OWASP Top 10 for LLM Applications (2025 version)