Security & governance · Glossary term

What is Jailbreak?

An adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are designed to prevent.

Why does Jailbreak matter?

Successful jailbreaks expose gaps between stated policy and actual behavior, and they can become more consequential when the model controls tools or protected data.

Jailbreak in practice

Derive test families from prohibited behaviors, vary format and interaction length, measure both refusal and harmful completion, and convert confirmed failures into versioned adversarial evals.

What is the common confusion about Jailbreak?

A jailbreak targets model or system behavioral restrictions. Prompt injection redirects instruction following, often toward an attacker's goal; one interaction can involve both.

Learn Jailbreak in the course

Start with

  • Capstone 82 — Jailbreak Taxonomy

    A safety harness without a taxonomy is a coin flip. Name the attack before you defend it. A model deployed without an attack model is a model defended against nothing in particular.

    Phase 19: Capstone Projects

Lessons that name Jailbreak in a title or section

  • ASCII Art and Visual Jailbreaks

    Jiang, Xu, Niu, Xiang, Ramasubramanian, Li, Poovendran, "ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs" (ACL 2024, arXiv:2402.11753).

    Phase 18: Ethics, Safety & Alignment

Taught in Phase 19: Capstone Projects.

Also covered in Phase 18: Ethics, Safety & Alignment.

  • Prompt InjectionAn attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or…
  • Red TeamingA structured adversarial testing process in which authorized testers seek failures using documented objectives, threat assumptions, cases,…
  • GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
  • Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…

Sources

More terms in Security & governance

Open the Security & governance list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.