Security & governance · Glossary term
What is Jailbreak?
An adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are designed to prevent.
Why does Jailbreak matter?
Successful jailbreaks expose gaps between stated policy and actual behavior, and they can become more consequential when the model controls tools or protected data.
Jailbreak in practice
Derive test families from prohibited behaviors, vary format and interaction length, measure both refusal and harmful completion, and convert confirmed failures into versioned adversarial evals.
What is the common confusion about Jailbreak?
A jailbreak targets model or system behavioral restrictions. Prompt injection redirects instruction following, often toward an attacker's goal; one interaction can involve both.
Learn Jailbreak in the course
Start with
- Capstone 82 — Jailbreak Taxonomy
A safety harness without a taxonomy is a coin flip. Name the attack before you defend it. A model deployed without an attack model is a model defended against nothing in particular.
Lessons that name Jailbreak in a title or section
- ASCII Art and Visual Jailbreaks
Jiang, Xu, Niu, Xiang, Ramasubramanian, Li, Poovendran, "ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs" (ACL 2024, arXiv:2402.11753).
Taught in Phase 19: Capstone Projects.
Also covered in Phase 18: Ethics, Safety & Alignment.
Related terms
- Prompt InjectionAn attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or…
- Red TeamingA structured adversarial testing process in which authorized testers seek failures using documented objectives, threat assumptions, cases,…
- GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
- Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
Sources
More terms in Security & governance
- AI Risk Assessment
- Audit Log
- Content Provenance
- Data Classification
- Data Exfiltration
- Data Lineage
- Data Minimization
- Datasheet for Datasets
- Defense in Depth
- Indirect Prompt Injection
- Membership Inference
- Provenance Attestation
- Purpose Limitation
- Red Teaming
- Separation of Duties
- Software Bill of Materials (SBOM)
- Threat Model
- Trust Boundary
- Zero Trust
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.