Evaluation & safety · Glossary term

What is Evaluation (Eval)?

A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and review procedures.

Also called Eval.

Why does Evaluation (Eval) matter?

You cannot improve reliability if success is only a subjective impression from a few demos.

Evaluation (Eval) in practice

Run the same customer-support scenarios before and after changing retrieval, score correctness and citation support, and inspect failures by category.

What is the common confusion about Evaluation (Eval)?

A benchmark score is one evaluation result, not a complete account of production quality.

Learn Evaluation (Eval) in the course

Start with

  • Evaluation & Testing LLM Applications

    You would never deploy a web app without tests. You would never ship a database migration without a rollback plan. But right now, most teams ship LLM applications by reading 10 outputs and saying…

    Phase 11: LLM Engineering

Lessons that name Evaluation (Eval) in a title or section

  • Model Evaluation

    A model is only as good as the way you measure it. Implement K-fold and stratified K-fold cross-validation from scratch and explain why stratification matters for imbalanced data.

    Phase 02: ML Fundamentals

  • LLM Evaluation — RAGAS, DeepEval, G-Eval

    Exact-match and F1 miss semantic equivalence. Human review does not scale. LLM-as-judge is the production answer — with enough calibration to trust the number.

    Phase 05: NLP: Foundations to Advanced

  • Long-Context Evaluation — NIAH, RULER, LongBench, MRCR

    Gemini 3 Pro advertises 10M tokens of context. At 1M tokens, 8-needle MRCR drops to 26.3%. Advertised ≠ usable. Long-context evaluation tells you the actual capacity of the model you are shipping on.

    Phase 05: NLP: Foundations to Advanced

  • Audio Evaluation — WER, MOS, UTMOS, MMAU, FAD, and the Open Leaderboards

    You cannot ship what you cannot measure. This lesson names the 2026 metrics for every audio task: ASR (WER, CER, RTFx), TTS (MOS, UTMOS, SECS, WER-on-ASR-round-trip), audio-language (MMAU,…

    Phase 06: Speech & Audio

  • Evaluation — FID, CLIP Score, Human Preference

    Every generative model leaderboard cites FID, CLIP score, and a win rate from a human-preference arena. Each number has a failure mode a determined researcher can game.

    Phase 08: Generative AI

  • Evaluation: Benchmarks, Evals, LM Harness

    Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. Every frontier lab games benchmarks. MMLU scores go up while models still can't reliably count the number of R's in…

    Phase 10: LLMs from Scratch

  • Skill Evals, Packaging, and Portability

    A skill is finished when its package survives linting, routes on the right requests, improves a measured task, stays inside policy, and degrades honestly on another host.

    Phase 13: Tools & Protocols

  • METR Time Horizons and External Capability Evaluation

    METR (ex-ARC Evals) is an independent 501(c)(3) since December 2023. Their Time Horizon 1.1 benchmark (January 2026) fits a logistic curve to task-success probability vs log(expert human completion…

    Phase 15: Autonomous Systems

Taught in Phase 11: LLM Engineering.

Also covered in Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 04: Computer Vision, Phase 05: NLP: Foundations to Advanced, Phase 06: Speech & Audio, Phase 08: Generative AI, Phase 09: Reinforcement Learning, Phase 10: LLMs from Scratch, Phase 12: Multimodal AI, Phase 13: Tools & Protocols, Phase 14: Agent Engineering, Phase 15: Autonomous Systems, Phase 16: Multi-Agent & Swarms, Phase 17: Infrastructure & Production, Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.

  • Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
  • LLM-as-a-JudgeUsing a language model to score, compare, classify, or critique another system's output against a rubric.
  • Cost per Successful TaskTotal system cost divided by the number of tasks that satisfy a defined success criterion, including retries, failed runs, tool use, and…
  • Regression TestA repeatable check that protects behavior known to work, especially after code, prompt, model, retrieval, or tool changes.
  • AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…
  • CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
  • Canary ReleaseA deployment strategy that exposes a new version to a limited slice of traffic or infrastructure before expanding the rollout.
  • Chain of Thought (CoT)Intermediate reasoning used to decompose a task before producing an answer. A prompt can request a visible rationale, while some systems…
  • GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
  • Late FusionProcessing modalities through separate encoders or predictors and combining their high-level representations, scores, or decisions near…
  • Loss FunctionAn objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce.
  • Model RouterA component that selects a model or provider for a request using requirements such as capability, latency, cost, context size, policy, and…
  • ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
  • PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
  • ROUGEA family of metrics that compares generated text with reference text using units such as n-gram overlap or longest common subsequence.
  • Shadow TrafficA copy of live request traffic sent to a candidate system for observation while the candidate response remains outside the primary user…
  • TraceA correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
  • Visual GroundingConnecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.

More terms in Evaluation & safety

Open the Evaluation & safety list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.