Evaluation & Quality
BERTScore
A score that compares AI-written text with an example answer using learned meanings of words in context.
Example
A team scores how semantically similar two phrasings appear to a pretrained AI system.
Why people use it
It compares meaning-related word patterns when checking generated writing against a reference.
What you'll hear
“The wording differs, but is the meaning similar?”
What this means for you
Use factual and task-specific checks alongside similarity scores.
Can you control it?
No
No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.
Common questions
- Can changing the example answer change the score?
- Yes. The comparison depends on the reference chosen, even when the AI's own answer stays the same.
- Can it miss a wrong factual detail?
- Yes. Two sentences can seem similar while disagreeing about a name, number or event.
- Is the score independent of the AI used to calculate it?
- No. The chosen comparison system can affect which similarities it recognizes.