Skip to content

Evaluation & Quality

BERTScore

A score that compares AI-written text with an example answer using learned meanings of words in context.

Example

A team scores how semantically similar two phrasings appear to a pretrained AI system.

Why people use it

It compares meaning-related word patterns when checking generated writing against a reference.

What you'll hear

“The wording differs, but is the meaning similar?”

What this means for you

Use factual and task-specific checks alongside similarity scores.

Can you control it?

No

No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.

Common questions

Can changing the example answer change the score?
Yes. The comparison depends on the reference chosen, even when the AI's own answer stays the same.
Can it miss a wrong factual detail?
Yes. Two sentences can seem similar while disagreeing about a name, number or event.
Is the score independent of the AI used to calculate it?
No. The chosen comparison system can affect which similarities it recognizes.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice