Evaluation & safety · Glossary term

What is LLM-as-a-Judge?

Using a language model to score, compare, classify, or critique another system's output against a rubric.

Why does LLM-as-a-Judge matter?

It can scale evaluation of qualities that are difficult to express as exact-match tests, such as clarity or instruction adherence.

LLM-as-a-Judge in practice

Give a separate evaluator model the task, candidate answer, reference evidence, and a structured rubric, then calibrate its scores against human-reviewed examples.

What is the common confusion about LLM-as-a-Judge?

A judge model is not ground truth. It can be biased by order, verbosity, style, prompt wording, or shared model failures.

Learn LLM-as-a-Judge in the course

Start with

  • Eval-Driven Agent Development

    Anthropic's guidance: "start with simple prompts, optimize them with comprehensive evaluation, and add multi-step agentic systems only when needed." Evaluation is not the last step.

    Phase 14: Agent Engineering

Taught in Phase 14: Agent Engineering.

  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
  • Verification GateA control point that blocks progress until defined evidence satisfies a correctness or quality criterion.
  • Precision & RecallPrecision asks how many flagged items were correct; recall asks how many relevant items were found.
  • Reviewer AgentAn agent assigned to inspect another agent's artifact or decision against explicit criteria and return findings or a verdict.
  • ROUGEA family of metrics that compares generated text with reference text using units such as n-gram overlap or longest common subsequence.

More terms in Evaluation & safety

Open the Evaluation & safety list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.