Evaluation & Quality
BLEU
Bilingual Evaluation Understudy
A score that compares word sequences in computer-written text with those in an example answer.
Example
A computer translation is compared with human examples by counting shared word sequences.
Why people use it
It offers a quick numerical comparison between generated wording and reference wording.
What you'll hear
“How much of the wording matches the checked translation?”
What this means for you
Include readers who know the language when judging translation quality.
Can you control it?
No
No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.
Common questions
- Does the example translation's quality matter?
- Yes. A poor or unusually worded reference can make the score less useful.
- Can a good translation get a low score?
- Yes. It may use valid wording that differs from the available references.
- Does a higher score prove better meaning?
- No. Matching word patterns does not by itself establish that the meaning is accurate.