Evaluation & safety · Glossary term

What are Precision & Recall?

Precision asks how many flagged items were correct; recall asks how many relevant items were found. When you change the decision threshold for one fixed scoring model, improving recall often lowers precision and vice versa. A better model can improve both. F1 is their harmonic mean.

What people say

“Two metrics for classification or retrieval quality.”

What is the common confusion about Precision & Recall?

The right threshold and metric depend on the cost of each error and the prevalence of the target class.

Learn Precision & Recall in the course

No lesson links to this term yet. Search the course catalog for it.

  • Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
  • Semantic SearchRetrieval that represents a query and candidates in an embedding space and ranks candidates using a vector-similarity function.
  • GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
  • CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
  • LLM-as-a-JudgeUsing a language model to score, compare, classify, or critique another system's output against a rubric.
  • Recall@KFor one query, Recall@K is `|relevant items intersecting the top k| / |relevant items|`.
  • ROUGEA family of metrics that compares generated text with reference text using units such as n-gram overlap or longest common subsequence.

More terms in Evaluation & safety

Open the Evaluation & safety list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.