Skip to content

Evaluation & Quality

ROUGE

Recall-Oriented Understudy for Gisting Evaluation

A family of scores comparing shared words or sequences between a summary and example summaries.

Example

A computer-written summary is compared with a human-written one.

Why people use it

It provides a quick check of how much wording a summary shares with a reference.

What you'll hear

“How much does this summary overlap with the checked one?”

What this means for you

Check the summary against the original source as well as comparing wording.

Can you control it?

No

No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.

Common questions

Does high ROUGE prove a summary is faithful?
No. Word overlap can coexist with incorrect or misleading claims.
Can a useful summary score poorly?
Yes. It may express the main points with different words from the reference.
Is there just one ROUGE score?
No. Different versions compare different kinds of word or sequence overlap.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice