Evaluation & safety · Glossary term
What is LLM-as-a-Judge?
Using a language model to score, compare, classify, or critique another system's output against a rubric.
Why does LLM-as-a-Judge matter?
It can scale evaluation of qualities that are difficult to express as exact-match tests, such as clarity or instruction adherence.
LLM-as-a-Judge in practice
Give a separate evaluator model the task, candidate answer, reference evidence, and a structured rubric, then calibrate its scores against human-reviewed examples.
What is the common confusion about LLM-as-a-Judge?
A judge model is not ground truth. It can be biased by order, verbosity, style, prompt wording, or shared model failures.
Learn LLM-as-a-Judge in the course
Start with
- Eval-Driven Agent Development
Anthropic's guidance: "start with simple prompts, optimize them with comprehensive evaluation, and add multi-step agentic systems only when needed." Evaluation is not the last step.
Taught in Phase 14: Agent Engineering.
Related terms
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
- Verification GateA control point that blocks progress until defined evidence satisfies a correctness or quality criterion.
- Precision & RecallPrecision asks how many flagged items were correct; recall asks how many relevant items were found.
- Reviewer AgentAn agent assigned to inspect another agent's artifact or decision against explicit criteria and return findings or a verdict.
- ROUGEA family of metrics that compares generated text with reference text using units such as n-gram overlap or longest common subsequence.
More terms in Evaluation & safety
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.