Evaluation & safety · Glossary term
What is Verification Gate?
A control point that blocks progress until defined evidence satisfies a correctness or quality criterion.
Why does Verification Gate matter?
It converts a model's claim of completion into an evidence-backed decision.
Verification Gate in practice
Prevent a coding task from completing until the patch applies, scoped tests pass, forbidden files remain unchanged, and required artifacts exist.
What is the common confusion about Verification Gate?
Verification checks whether evidence meets criteria. Approval grants authority to proceed, even when the evidence is already known.
Learn Verification Gate in the course
Start with
- Verification Gates
The agent does not get to mark its own work as done. A verification gate reads the scope contract, the feedback log, the rule report, and the diff, and answers a single question: is this task…
Lessons that name Verification Gate in a title or section
- Capstone Lesson 25: Verification Gates and the Observation Budget
An agent harness without a verification layer is a wish in a trenchcoat. This lesson builds the deterministic gate chain that decides whether a tool call is allowed to fire, how much of its output…
- Reviewer Agent: Separate Builder from Marker
The agent that wrote the code cannot grade it. A reviewer is a second loop with a different system prompt, a different goal, and read-only access to everything the builder produced.
Taught in Phase 14: Agent Engineering.
Also covered in Phase 19: Capstone Projects.
Related terms
- Approval GateA control point that blocks a consequential action until an authorized person or policy grants permission.
- Regression TestA repeatable check that protects behavior known to work, especially after code, prompt, model, retrieval, or tool changes.
- Scope ContractA concrete agreement that defines a task's goal, allowed and forbidden surfaces, expected artifacts, verification requirements, and…
- Structured OutputModel output constrained or validated against a machine-readable schema so application code can consume fields without parsing free-form…
- Agent HarnessThe runtime around a model that assembles context, exposes tools, manages state, enforces limits, records traces, and decides when the…
- Canary ReleaseA deployment strategy that exposes a new version to a limited slice of traffic or infrastructure before expanding the rollout.
- Chain of Thought (CoT)Intermediate reasoning used to decompose a task before producing an answer. A prompt can request a visible rationale, while some systems…
- Cost per Successful TaskTotal system cost divided by the number of tasks that satisfy a defined success criterion, including retries, failed runs, tool use, and…
- Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
- Flaky TestA test that can pass and fail across equivalent runs without a relevant change to the code or intended test environment.
- GroundingConnecting a generated answer or action to evidence, state, or observations that the system can identify and check.
- HallucinationGenerated content that is false, unsupported by the available evidence, or inconsistent with the task's source of truth.
- Human-in-the-Loop (HITL)A workflow design in which a person supplies judgment, correction, approval, or escalation at defined points in an AI-driven process.
- LLM-as-a-JudgeUsing a language model to score, compare, classify, or critique another system's output against a rubric.
- PlanningConstructing, selecting, or revising a sequence of actions and dependencies intended to move from the current state to a goal.
- Provenance AttestationAuthenticated, machine-readable metadata that binds an artifact to claims about how, where, when, and from which inputs it was produced.
- Reproducible BuildA build whose declared source, environment, and instructions can be independently rerun to produce bit-for-bit identical specified…
- Reviewer AgentAn agent assigned to inspect another agent's artifact or decision against explicit criteria and return findings or a verdict.
- Termination ConditionAn explicit rule that ends or pauses an agent run when it succeeds, fails, exhausts a budget, reaches a safe boundary, or requires…
- Test OracleThe mechanism, specification, reference, invariant, or human judgment used to decide whether observed program behavior is correct.
More terms in Evaluation & safety
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.