ROUGE
(Recall-Oriented Understudy for Gisting Evaluation) is a set of metrics used to evaluate automatically generated text by comparing it with one or more human-written reference texts.
What is ROUGE?
ROUGE measures the overlap between generated and reference text using units such as words, word sequences, or sequences of characters. Common variants include ROUGE-1, which compares individual words, ROUGE-2, which compares two-word sequences, and ROUGE-L, which uses the longest common subsequence.
Why is ROUGE Important?
ROUGE provides a quantitative way to evaluate text-generation systems, particularly when assessing how much relevant content from a reference appears in a generated output. However, overlap-based scores may not fully capture meaning, factual accuracy, or linguistic quality.
Common use cases
ROUGE is commonly used to evaluate text summarization, machine translation research, question answering, and other natural language generation systems.