Skip to content

Evaluation & Quality

BLEU

Bilingual Evaluation Understudy

A score that compares word sequences in computer-written text with those in an example answer.

Example

A computer translation is compared with human examples by counting shared word sequences.

Why people use it

It offers a quick numerical comparison between generated wording and reference wording.

What you'll hear

“How much of the wording matches the checked translation?”

What this means for you

Include readers who know the language when judging translation quality.

Can you control it?

No

No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.

Common questions

Does the example translation's quality matter?
Yes. A poor or unusually worded reference can make the score less useful.
Can a good translation get a low score?
Yes. It may use valid wording that differs from the available references.
Does a higher score prove better meaning?
No. Matching word patterns does not by itself establish that the meaning is accurate.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice