Evaluation & safety · Glossary term
What are Precision & Recall?
Precision asks how many flagged items were correct; recall asks how many relevant items were found. When you change the decision threshold for one fixed scoring model, improving recall often lowers precision and vice versa. A better model can improve both. F1 is their harmonic mean.
“Two metrics for classification or retrieval quality.”
What is the common confusion about Precision & Recall?
The right threshold and metric depend on the cost of each error and the prevalence of the target class.
Learn Precision & Recall in the course
No lesson links to this term yet. Search the course catalog for it.
Related terms
- Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
- Semantic SearchRetrieval that represents a query and candidates in an embedding space and ranks candidates using a vector-similarity function.
- GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
- CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
- LLM-as-a-JudgeUsing a language model to score, compare, classify, or critique another system's output against a rubric.
- Recall@KFor one query, Recall@K is `|relevant items intersecting the top k| / |relevant items|`.
- ROUGEA family of metrics that compares generated text with reference text using units such as n-gram overlap or longest common subsequence.
More terms in Evaluation & safety
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.