Evaluation & safety · Glossary term
What are Guardrails?
System controls that constrain inputs, tool use, outputs, permissions, and escalation. They can include schemas, policy checks, classifiers, allowlists, sandboxing, approvals, and post-action verification.
“Safety filters around a model.”
Why do Guardrails matter?
No single filter covers all failure modes, so controls should be layered according to risk.
What is the common confusion about Guardrails?
Guardrails reduce risk; they do not prove that an AI system is safe.
Learn Guardrails in the course
Start with
- Guardrails, Safety & Content Filtering
Your LLM application will be attacked. Not might. Will. The first prompt injection attempt against your production system will come within 48 hours of launch.
Lessons that name Guardrails in a title or section
- OpenAI Agents SDK: Handoffs, Guardrails, Tracing
OpenAI Agents SDK is the lightweight multi-agent framework built on the Responses API. Five primitives: Agent, Handoff, Guardrail, Session, Tracing. Handoffs are tools named transferto .
- Security — Secrets, API Key Rotation, Audit Logs, Guardrails
Eliminate secret sprawl via centralized vaults (HashiCorp Vault, AWS Secrets Manager, Azure Key Vault). Never store credentials in config files, env files in VCS, spreadsheets.
- Building a Production LLM Application
You have built prompts, embeddings, RAG pipelines, function calling, caching layers, and guardrails. Separately. In isolation. Like practicing guitar scales without ever playing a song.
- LLM Routing Layer — LiteLLM, OpenRouter, Portkey
Provider lock-in is expensive. Different tool-calling workloads suit different models. Routing gateways give one API surface, retries, failover, cost tracking, and guardrails.
- Self-Refine and CRITIC: Iterative Output Improvement
Self-Refine (Madaan et al., 2023) uses one LLM in three roles — generate, feedback, refine — in a loop. Average gain: +20 absolute on 7 tasks.
- Agent Instructions as Executable Constraints
Instructions written as prose are wishes. Instructions written as constraints are tests. The workbench turns each rule into something an agent can check at runtime and a reviewer can verify after…
- Llama Guard and Input/Output Classification
Llama Guard 3 (Meta, Llama-3.1-8B base, fine-tuned for content safety) classifies both LLM inputs and outputs against an MLCommons 13-hazard taxonomy across 8 languages.
- Chaos Engineering for LLM Production
Chaos engineering for LLMs is its own discipline in 2026. Prerequisites before running experiments in production: defined SLI/SLO, trace+metric+log observability, automated rollback, runbooks,…
Taught in Phase 11: LLM Engineering.
Also covered in Phase 13: Tools & Protocols, Phase 14: Agent Engineering, Phase 15: Autonomous Systems and Phase 17: Infrastructure & Production.
Related terms
- Least PrivilegeGiving a model, agent, tool, or user only the permissions required for the current task, for only as long as those permissions are needed.
- Approval GateA control point that blocks a consequential action until an authorized person or policy grants permission.
- SandboxAn isolated execution environment that restricts an agent's access to files, processes, network destinations, credentials, and host…
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- AI Risk AssessmentA documented analysis of how an AI system can affect people, organizations, and environments, including context, hazards, likelihood,…
- AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…
- Defense in DepthUsing independent preventive, detective, and corrective controls at several system boundaries so one failed control does not determine the…
- Human-in-the-Loop (HITL)A workflow design in which a person supplies judgment, correction, approval, or escalation at defined points in an AI-driven process.
- JailbreakAn adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are…
- Precision & RecallPrecision asks how many flagged items were correct; recall asks how many relevant items were found.
- Red TeamingA structured adversarial testing process in which authorized testers seek failures using documented objectives, threat assumptions, cases,…
- System PromptA provider-defined instruction message or configuration supplied by the application to establish behavior and constraints within that…
More terms in Evaluation & safety
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.