Skip to content

Evaluation & Quality

Agent evaluation

Checking how well an AI helper carries out tasks, including its steps and final results.

Example

An evaluator checks whether a booking agent actually created the correct reservation.

Why people use it

It checks whether an AI helper actually finishes tasks correctly.

What you'll hear

“Did the agent complete the booking or just say it did?”

What this means for you

Check what changed in the basic system, not only what the agent reports.

Can you control it?

Developer-only

The people building or running the AI choose this setup. An everyday user generally needs their help to change how this part works.

Common questions

Is a convincing final message enough to pass?
No. The real outcome and any unauthorized actions also matter.
Should failed attempts count?
Yes. Abandoned tasks, repeated attempts and mistakes matter when judging how dependable the agent is.
Can an agent pass a test but struggle at work?
Yes. Real accounts, unexpected messages and changing websites can present situations the test never included.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice