Evaluation & safety · Glossary term

What is Eval Set?

A versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined capability or risk.

Also called Evaluation set.

Why does Eval Set matter?

A repeatable set turns vague quality claims into comparable evidence and catches regressions after prompts, models, tools, or retrieval change.

Eval Set in practice

Keep representative support questions, adversarial instructions, expected citations, and failure labels in a reviewed dataset that is separate from development examples.

What is the common confusion about Eval Set?

A development eval guides iteration, a final held-out test estimates performance after choices are fixed, and a standardized benchmark supports comparison under a shared protocol. Repeated tuning against any held-out set leaks test information and inflates results.

Learn Eval Set in the course

Start with

  • Eval-Driven Agent Development

    Anthropic's guidance: "start with simple prompts, optimize them with comprehensive evaluation, and add multi-step agentic systems only when needed." Evaluation is not the last step.

    Phase 14: Agent Engineering

Taught in Phase 14: Agent Engineering.

  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • Regression TestA repeatable check that protects behavior known to work, especially after code, prompt, model, retrieval, or tool changes.
  • LLM-as-a-JudgeUsing a language model to score, compare, classify, or critique another system's output against a rubric.
  • Verification GateA control point that blocks progress until defined evidence satisfies a correctness or quality criterion.
  • Benchmark ContaminationOverlap or information leakage between evaluation examples and data used to pretrain, tune, prompt, select, or otherwise improve the…
  • Data AugmentationCreating modified examples, such as transformed images, perturbed audio, or paraphrased text, to increase training diversity without…
  • Data LeakageUnintended use of information during training or feature construction that would not be available at the real prediction point or belongs…
  • Dataset SplitA documented partition of examples into separate subsets for fitting, development decisions, and final evaluation.
  • Distribution ShiftA difference between the data distribution used to build or evaluate a system and the distribution it encounters after deployment.
  • EpochOne traversal of the defined training dataset. In distributed or sampled training, the exact implementation of an epoch depends on the…
  • Exact Match (EM)A metric that counts an output as correct only when its normalized representation exactly equals an accepted reference answer.
  • JailbreakAn adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are…
  • Lost in the MiddleA long-context failure pattern in which model performance changes with evidence position and can degrade when relevant information sits…
  • Membership InferenceAn attack that estimates whether a particular record or example was included in a model's training data by observing model outputs or…
  • Model CardA structured report describing a model's intended uses, evaluation conditions, performance characteristics, limitations, and relevant…
  • OverfittingA generalization gap in which performance on training data is substantially better than performance on representative unseen data.
  • Pass@kAcross a task set, the fraction of tasks for which at least one of k sampled candidates passes a defined correctness test.
  • Precision & RecallPrecision asks how many flagged items were correct; recall asks how many relevant items were found.
  • Prompt SensitivityVariation in model output or measured performance caused by changes to prompt wording, order, formatting, or examples that preserve the…
  • Recall@KFor one query, Recall@K is `|relevant items intersecting the top k| / |relevant items|`.
  • Red TeamingA structured adversarial testing process in which authorized testers seek failures using documented objectives, threat assumptions, cases,…
  • Test OracleThe mechanism, specification, reference, invariant, or human judgment used to decide whether observed program behavior is correct.

More terms in Evaluation & safety

Open the Evaluation & safety list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.