Evaluation & safety · Glossary term

What are Guardrails?

System controls that constrain inputs, tool use, outputs, permissions, and escalation. They can include schemas, policy checks, classifiers, allowlists, sandboxing, approvals, and post-action verification.

What people say

“Safety filters around a model.”

Why do Guardrails matter?

No single filter covers all failure modes, so controls should be layered according to risk.

What is the common confusion about Guardrails?

Guardrails reduce risk; they do not prove that an AI system is safe.

Learn Guardrails in the course

Start with

  • Guardrails, Safety & Content Filtering

    Your LLM application will be attacked. Not might. Will. The first prompt injection attempt against your production system will come within 48 hours of launch.

    Phase 11: LLM Engineering

Lessons that name Guardrails in a title or section

  • OpenAI Agents SDK: Handoffs, Guardrails, Tracing

    OpenAI Agents SDK is the lightweight multi-agent framework built on the Responses API. Five primitives: Agent, Handoff, Guardrail, Session, Tracing. Handoffs are tools named transferto .

    Phase 14: Agent Engineering

  • Security — Secrets, API Key Rotation, Audit Logs, Guardrails

    Eliminate secret sprawl via centralized vaults (HashiCorp Vault, AWS Secrets Manager, Azure Key Vault). Never store credentials in config files, env files in VCS, spreadsheets.

    Phase 17: Infrastructure & Production

  • Building a Production LLM Application

    You have built prompts, embeddings, RAG pipelines, function calling, caching layers, and guardrails. Separately. In isolation. Like practicing guitar scales without ever playing a song.

    Phase 11: LLM Engineering

  • LLM Routing Layer — LiteLLM, OpenRouter, Portkey

    Provider lock-in is expensive. Different tool-calling workloads suit different models. Routing gateways give one API surface, retries, failover, cost tracking, and guardrails.

    Phase 13: Tools & Protocols

  • Self-Refine and CRITIC: Iterative Output Improvement

    Self-Refine (Madaan et al., 2023) uses one LLM in three roles — generate, feedback, refine — in a loop. Average gain: +20 absolute on 7 tasks.

    Phase 14: Agent Engineering

  • Agent Instructions as Executable Constraints

    Instructions written as prose are wishes. Instructions written as constraints are tests. The workbench turns each rule into something an agent can check at runtime and a reviewer can verify after…

    Phase 14: Agent Engineering

  • Llama Guard and Input/Output Classification

    Llama Guard 3 (Meta, Llama-3.1-8B base, fine-tuned for content safety) classifies both LLM inputs and outputs against an MLCommons 13-hazard taxonomy across 8 languages.

    Phase 15: Autonomous Systems

  • Chaos Engineering for LLM Production

    Chaos engineering for LLMs is its own discipline in 2026. Prerequisites before running experiments in production: defined SLI/SLO, trace+metric+log observability, automated rollback, runbooks,…

    Phase 17: Infrastructure & Production

Taught in Phase 11: LLM Engineering.

Also covered in Phase 13: Tools & Protocols, Phase 14: Agent Engineering, Phase 15: Autonomous Systems and Phase 17: Infrastructure & Production.

  • Least PrivilegeGiving a model, agent, tool, or user only the permissions required for the current task, for only as long as those permissions are needed.
  • Approval GateA control point that blocks a consequential action until an authorized person or policy grants permission.
  • SandboxAn isolated execution environment that restricts an agent's access to files, processes, network destinations, credentials, and host…
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • AI Risk AssessmentA documented analysis of how an AI system can affect people, organizations, and environments, including context, hazards, likelihood,…
  • AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…
  • Defense in DepthUsing independent preventive, detective, and corrective controls at several system boundaries so one failed control does not determine the…
  • Human-in-the-Loop (HITL)A workflow design in which a person supplies judgment, correction, approval, or escalation at defined points in an AI-driven process.
  • JailbreakAn adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are…
  • Precision & RecallPrecision asks how many flagged items were correct; recall asks how many relevant items were found.
  • Red TeamingA structured adversarial testing process in which authorized testers seek failures using documented objectives, threat assumptions, cases,…
  • System PromptA provider-defined instruction message or configuration supplied by the application to establish behavior and constraints within that…

More terms in Evaluation & safety

Open the Evaluation & safety list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.