Evaluation & safety · Glossary term

What is Calibration?

The agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.

Why does Calibration matter?

A system can be accurate on average yet dangerously overconfident on the cases where people rely on its score.

Calibration in practice

Bucket predictions by confidence, compare confidence with empirical accuracy, and recalibrate or abstain when the gap is unacceptable.

What is the common confusion about Calibration?

Calibration measures confidence reliability, not overall accuracy, factuality, or reasoning quality.

Learn Calibration in the course

Lessons that name Calibration in a title or section

  • Perplexity and Calibration

    If your model says 90 percent confident on a thousand answers and gets six hundred right, it is not well calibrated. Calibration is half of trustworthy eval.

    Phase 19: Capstone Projects

  • Production Quantization — AWQ, GPTQ, GGUF K-quants, FP8, MXFP4/NVFP4

    Quantization format is not a universal choice — it is a function of hardware, serving engine, and workload. GGUF Q4KM or Q5KM owns CPU and edge, delivered through llama.cpp and Ollama.

    Phase 17: Infrastructure & Production

  • Sycophancy as RLHF Amplification

    Sycophancy is not a bug in the data — it is a property of the loss. Shapira et al. (arXiv:2602.01002, Feb 2026) give a formal two-stage mechanism: sycophantic completions are over-represented among…

    Phase 18: Ethics, Safety & Alignment

  • EchoLeak and the Emergence of CVEs for AI

    CVE-2025-32711 "EchoLeak" (CVSS 9.3) was the first publicly documented zero-click prompt injection in a production LLM system (Microsoft 365 Copilot).

    Phase 18: Ethics, Safety & Alignment

Covered in Phase 17: Infrastructure & Production, Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.

  • SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • Precision & RecallPrecision asks how many flagged items were correct; recall asks how many relevant items were found.
  • LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…

Sources

More terms in Evaluation & safety

Open the Evaluation & safety list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.