Agent evals
Why it matters
An agent can demo well and still fail on a tenth of real requests, in ways nobody notices until a customer hits the gap. The same agent can also answer differently each time it runs, so a few manual checks prove little. Evals turn that risk into a number. When a prompt is rewritten or a model is swapped, the score shows whether quality rose or fell, instead of leaving the change as a guess made on production traffic. Without evals, every change to the agent is a roll of the dice.
How to apply it
- Collect real tasks the agent will handle, for example last month's support tickets with the correct replies.
- Include hard and awkward cases, not only the easy wins.
- Write the pass rule for each task before running anything.
- Run the whole set after every change to the prompt, the model or the tools the agent can call.
- Run it more than once, since results can vary from run to run.
- Add every real failure to the set so it cannot return unnoticed.
What it is
An eval, short for evaluation, is a test for an AI agent. It has a set of example tasks, a description of what a good result looks like and a way of scoring each answer. Scoring can be an exact match, a rule such as "the refund amount must equal the order total", a person's review or another model acting as judge. The result is a number, such as 46 of 50 tasks passed, which can be tracked over time.
Common mistakes
- Checking quality once before launch and never again.
- Using a model as judge without sampling its scores against a person's.