Skip to content

Evaluation & Quality

Brier score

A score measuring how far predicted chances are from what actually happens, with larger mistakes counting more.

Example

A binary forecast of 0.8 is compared with an outcome recorded as one.

Why people use it

It checks how close stated chances are to the outcomes that occurred.

What you'll hear

“How good were those probability forecasts?”

What this means for you

State the task and scoring convention when comparing results.

Can you control it?

No

No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.

Common questions

Can two systems get the same score in different ways?
Yes. One overall number can hide different patterns of confident mistakes and useful distinctions between cases.
Is a lower score generally better?
Yes, when the same task and scoring rules are used.
Does it judge only whether the most likely choice was right?
No. The stated probabilities matter, not just the final category chosen.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice