Evaluation & safety · Glossary term
What is Evaluation (Eval)?
A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and review procedures.
Also called Eval.
Why does Evaluation (Eval) matter?
You cannot improve reliability if success is only a subjective impression from a few demos.
Evaluation (Eval) in practice
Run the same customer-support scenarios before and after changing retrieval, score correctness and citation support, and inspect failures by category.
What is the common confusion about Evaluation (Eval)?
A benchmark score is one evaluation result, not a complete account of production quality.
Learn Evaluation (Eval) in the course
Start with
- Evaluation & Testing LLM Applications
You would never deploy a web app without tests. You would never ship a database migration without a rollback plan. But right now, most teams ship LLM applications by reading 10 outputs and saying…
Lessons that name Evaluation (Eval) in a title or section
- Model Evaluation
A model is only as good as the way you measure it. Implement K-fold and stratified K-fold cross-validation from scratch and explain why stratification matters for imbalanced data.
- LLM Evaluation — RAGAS, DeepEval, G-Eval
Exact-match and F1 miss semantic equivalence. Human review does not scale. LLM-as-judge is the production answer — with enough calibration to trust the number.
- Long-Context Evaluation — NIAH, RULER, LongBench, MRCR
Gemini 3 Pro advertises 10M tokens of context. At 1M tokens, 8-needle MRCR drops to 26.3%. Advertised ≠ usable. Long-context evaluation tells you the actual capacity of the model you are shipping on.
- Audio Evaluation — WER, MOS, UTMOS, MMAU, FAD, and the Open Leaderboards
You cannot ship what you cannot measure. This lesson names the 2026 metrics for every audio task: ASR (WER, CER, RTFx), TTS (MOS, UTMOS, SECS, WER-on-ASR-round-trip), audio-language (MMAU,…
- Evaluation — FID, CLIP Score, Human Preference
Every generative model leaderboard cites FID, CLIP score, and a win rate from a human-preference arena. Each number has a failure mode a determined researcher can game.
- Evaluation: Benchmarks, Evals, LM Harness
Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. Every frontier lab games benchmarks. MMLU scores go up while models still can't reliably count the number of R's in…
- Skill Evals, Packaging, and Portability
A skill is finished when its package survives linting, routes on the right requests, improves a measured task, stays inside policy, and degrades honestly on another host.
- METR Time Horizons and External Capability Evaluation
METR (ex-ARC Evals) is an independent 501(c)(3) since December 2023. Their Time Horizon 1.1 benchmark (January 2026) fits a logistic curve to task-success probability vs log(expert human completion…
Taught in Phase 11: LLM Engineering.
Also covered in Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 04: Computer Vision, Phase 05: NLP: Foundations to Advanced, Phase 06: Speech & Audio, Phase 08: Generative AI, Phase 09: Reinforcement Learning, Phase 10: LLMs from Scratch, Phase 12: Multimodal AI, Phase 13: Tools & Protocols, Phase 14: Agent Engineering, Phase 15: Autonomous Systems, Phase 16: Multi-Agent & Swarms, Phase 17: Infrastructure & Production, Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.
Related terms
- Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
- LLM-as-a-JudgeUsing a language model to score, compare, classify, or critique another system's output against a rubric.
- Cost per Successful TaskTotal system cost divided by the number of tasks that satisfy a defined success criterion, including retries, failed runs, tool use, and…
- Regression TestA repeatable check that protects behavior known to work, especially after code, prompt, model, retrieval, or tool changes.
- AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…
- CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
- Canary ReleaseA deployment strategy that exposes a new version to a limited slice of traffic or infrastructure before expanding the rollout.
- Chain of Thought (CoT)Intermediate reasoning used to decompose a task before producing an answer. A prompt can request a visible rationale, while some systems…
- GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
- Late FusionProcessing modalities through separate encoders or predictors and combining their high-level representations, scores, or decisions near…
- Loss FunctionAn objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce.
- Model RouterA component that selects a model or provider for a request using requirements such as capability, latency, cost, context size, policy, and…
- ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
- PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
- ROUGEA family of metrics that compares generated text with reference text using units such as n-gram overlap or longest common subsequence.
- Shadow TrafficA copy of live request traffic sent to a candidate system for observation while the candidate response remains outside the primary user…
- TraceA correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
- Visual GroundingConnecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.
More terms in Evaluation & safety
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.