Evaluation & safety · Glossary term

What is Prompt Injection?

An attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or take actions outside the user's goal. The content can arrive directly from a user or indirectly through retrieved pages, files, messages, or tool output.

What people say

“An adversarial instruction that redirects a model.”

Why does Prompt Injection matter?

Models process instructions and data through the same language channel, so input filtering alone cannot reliably separate every malicious instruction from legitimate content.

Prompt Injection in practice

Treat external content as untrusted, isolate it from authority-bearing instructions, minimize tool permissions, require approval for consequential writes, and verify outputs and actions.

What is the common confusion about Prompt Injection?

Prompt injection is not technically the same mechanism as SQL injection, and a stronger system prompt is not a complete defense.

Learn Prompt Injection in the course

Start with

  • Prompt Injection and the PVE Defense

    Greshake et al. (AISec 2023) established indirect prompt injection as the defining agent security problem. Attacker plants instructions in data the agent retrieves; on ingest, those instructions…

    Phase 14: Agent Engineering

Lessons that name Prompt Injection in a title or section

  • Indirect Prompt Injection — Production Attack Surface

    Indirect prompt injection (IPI) embeds instructions inside external content — a web page, an email, a shared document, a support ticket — consumed by an agentic system without explicit user action.

    Phase 18: Ethics, Safety & Alignment

  • Capstone 83 — Prompt Injection Detector

    A detector is a function from prompt to confidence and category. Anything else is a vibe. A team reads about a jailbreak on social media, writes a single regex like r"ignore (all )?previous", ships…

    Phase 19: Capstone Projects

Taught in Phase 14: Agent Engineering.

Also covered in Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.

  • Least PrivilegeGiving a model, agent, tool, or user only the permissions required for the current task, for only as long as those permissions are needed.
  • SandboxAn isolated execution environment that restricts an agent's access to files, processes, network destinations, credentials, and host…
  • Approval GateA control point that blocks a consequential action until an authorized person or policy grants permission.
  • Tool ContractThe complete agreement for a tool boundary: purpose, typed inputs, outputs, validation, permissions, side effects, errors, timeouts,…
  • Indirect Prompt InjectionA prompt-injection attack delivered through content the system retrieves or observes, such as a webpage, document, email, image text, or…
  • Instruction HierarchyA rule set for resolving conflicts among instructions from sources with different authority, such as application policy, users, and…
  • JailbreakAn adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are…
  • Red TeamingA structured adversarial testing process in which authorized testers seek failures using documented objectives, threat assumptions, cases,…
  • System PromptA provider-defined instruction message or configuration supplied by the application to establish behavior and constraints within that…
  • Threat ModelA documented account of protected assets, trust boundaries, potential adversaries, assumed capabilities, attack paths, impacts, and…

Sources

More terms in Evaluation & safety

Open the Evaluation & safety list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.