Evaluation & safety · Glossary term
What is Calibration?
The agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
Why does Calibration matter?
A system can be accurate on average yet dangerously overconfident on the cases where people rely on its score.
Calibration in practice
Bucket predictions by confidence, compare confidence with empirical accuracy, and recalibrate or abstain when the gap is unacceptable.
What is the common confusion about Calibration?
Calibration measures confidence reliability, not overall accuracy, factuality, or reasoning quality.
Learn Calibration in the course
Lessons that name Calibration in a title or section
- Perplexity and Calibration
If your model says 90 percent confident on a thousand answers and gets six hundred right, it is not well calibrated. Calibration is half of trustworthy eval.
- Production Quantization — AWQ, GPTQ, GGUF K-quants, FP8, MXFP4/NVFP4
Quantization format is not a universal choice — it is a function of hardware, serving engine, and workload. GGUF Q4KM or Q5KM owns CPU and edge, delivered through llama.cpp and Ollama.
- Sycophancy as RLHF Amplification
Sycophancy is not a bug in the data — it is a property of the loss. Shapira et al. (arXiv:2602.01002, Feb 2026) give a formal two-stage mechanism: sycophantic completions are over-represented among…
- EchoLeak and the Emergence of CVEs for AI
CVE-2025-32711 "EchoLeak" (CVSS 9.3) was the first publicly documented zero-click prompt injection in a production LLM system (Microsoft 365 Copilot).
Covered in Phase 17: Infrastructure & Production, Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.
Related terms
- SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- Precision & RecallPrecision asks how many flagged items were correct; recall asks how many relevant items were found.
- LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…
Sources
More terms in Evaluation & safety
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.