Evaluation & safety · Glossary term
What is ROUGE?
A family of metrics that compares generated text with reference text using units such as n-gram overlap or longest common subsequence.
“A reference-overlap metric often used for summaries.”
What is the common confusion about ROUGE?
Surface overlap can miss semantic equivalence and can reward copied wording without proving factual quality.
Learn ROUGE in the course
Lessons that name ROUGE in a title or section
- Text Summarization
Extractive systems tell you what the document said. Abstractive systems tell you what the author meant. Different tasks, different pitfalls. A 2,000-word news article lands in your feed.
Covered in Phase 05: NLP: Foundations to Advanced.
Related terms
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- Precision & RecallPrecision asks how many flagged items were correct; recall asks how many relevant items were found.
- LLM-as-a-JudgeUsing a language model to score, compare, classify, or critique another system's output against a rubric.
- Exact Match (EM)A metric that counts an output as correct only when its normalized representation exactly equals an accepted reference answer.
More terms in Evaluation & safety
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.