Skip to content

Evaluation & Quality

LLM-as-a-judge

Using a language model to score or compare outputs from another AI system.

Example

A model grades draft answers against a rubric before human spot-checking.

Why people use it

Teams use “LLM-as-a-judge” when they need to measure quality with evidence instead of impressions.

What you'll hear

“What does LLM-as-a-judge tell us about whether the system is actually working?”

Related terms