Agent evals

Definition
Agent evals are the tests that score whether an AI agent's outputs are actually right, against known-good answers or real outcomes, rather than trusting that they look right.

Why it matters

An agent can demo well and still fail on a tenth of real requests, in ways nobody notices until a customer hits the gap. The same agent can also answer differently each time it runs, so a few manual checks prove little. Evals turn that risk into a number. When a prompt is rewritten or a model is swapped, the score shows whether quality rose or fell, instead of leaving the change as a guess made on production traffic. Without evals, every change to the agent is a roll of the dice.

How to apply it

  • Collect real tasks the agent will handle, for example last month's support tickets with the correct replies.
  • Include hard and awkward cases, not only the easy wins.
  • Write the pass rule for each task before running anything.
  • Run the whole set after every change to the prompt, the model or the tools the agent can call.
  • Run it more than once, since results can vary from run to run.
  • Add every real failure to the set so it cannot return unnoticed.

What it is

An eval, short for evaluation, is a test for an AI agent. It has a set of example tasks, a description of what a good result looks like and a way of scoring each answer. Scoring can be an exact match, a rule such as "the refund amount must equal the order total", a person's review or another model acting as judge. The result is a number, such as 46 of 50 tasks passed, which can be tracked over time.

Common mistakes

  • Checking quality once before launch and never again.
  • Using a model as judge without sampling its scores against a person's.
Worked example

Suppose a Series A software company runs an AI support agent, built in Botpress, that answers billing questions. Before a prompt change, the team pulls 50 real tickets from last month, each with a correct reply written by a senior agent. Every answer is scored against a rule: any refund quoted must equal the order total on record. The current version passes 43 of 50. After a rewrite of the prompt, the score rises to 46 of 50. Because answers vary between runs, the team repeats the set three times and releases only when the score holds across all three. Each new real failure is added to the set, so it cannot return unnoticed. A demo that looked right is not treated as proof.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    Regression

    A change that breaks something that used to work.

  2. Article

    Test coverage

    The equivalent measure for ordinary code.

  3. Article

    Hallucination

    The confident wrong answer evals try to catch.