Prompting & context · Glossary term
What is Instruction Hierarchy?
A rule set for resolving conflicts among instructions from sources with different authority, such as application policy, users, and untrusted retrieved content.
Why does Instruction Hierarchy matter?
Agent systems mix trusted goals with external text, so the model and harness need a defined response when lower-authority content conflicts with higher-authority constraints.
Instruction Hierarchy in practice
Label untrusted tool output as data, preserve higher-priority constraints outside that content, and test direct and indirect conflict cases.
What is the common confusion about Instruction Hierarchy?
An instruction hierarchy can improve behavior but is not a security boundary; least privilege and approval controls still limit consequences.
Learn Instruction Hierarchy in the course
No lesson links to this term yet. Search the course catalog for it.
Related terms
- System PromptA provider-defined instruction message or configuration supplied by the application to establish behavior and constraints within that…
- Prompt InjectionAn attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or…
- Least PrivilegeGiving a model, agent, tool, or user only the permissions required for the current task, for only as long as those permissions are needed.
- Tool ContractThe complete agreement for a tool boundary: purpose, typed inputs, outputs, validation, permissions, side effects, errors, timeouts,…
- Indirect Prompt InjectionA prompt-injection attack delivered through content the system retrieves or observes, such as a webpage, document, email, image text, or…
- Instruction FollowingA model capability to map natural-language directions and supplied context to behavior that satisfies the stated task and constraints.
Sources
More terms in Prompting & context
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.