Evaluation & Quality
Agent evaluation
Checking how well an AI helper carries out tasks, including its steps and final results.
Example
An evaluator checks whether a booking agent actually created the correct reservation.
Why people use it
It checks whether an AI helper actually finishes tasks correctly.
What you'll hear
“Did the agent complete the booking or just say it did?”
What this means for you
Check what changed in the basic system, not only what the agent reports.
Can you control it?
Developer-only
The people building or running the AI choose this setup. An everyday user generally needs their help to change how this part works.
Common questions
- Is a convincing final message enough to pass?
- No. The real outcome and any unauthorized actions also matter.
- Should failed attempts count?
- Yes. Abandoned tasks, repeated attempts and mistakes matter when judging how dependable the agent is.
- Can an agent pass a test but struggle at work?
- Yes. Real accounts, unexpected messages and changing websites can present situations the test never included.