Traditional tests assert exact outputs. AI features rarely produce the same output twice, so teams either skip testing or rely on "it looked fine when I tried it". Both lead to silent regressions when you change a prompt, a model or a retrieval setting.

Evals are the answer: repeatable checks that score AI behaviour on a fixed set of cases.

Start with a golden dataset

Collect 30–100 real examples of the task, each with the input and what a good answer must contain. Pull them from:

  • Real user requests (anonymised).
  • Known failure cases and bug reports.
  • Edge cases: empty input, very long input, other languages, adversarial input.

Small and representative beats large and synthetic.

Choose the right scorer for each check

Use the cheapest scorer that works.

  1. Deterministic assertions – valid JSON, required fields present, length limits, no banned phrases, correct tool called.
  2. Reference comparison – does the answer contain the key facts from the expected answer?
  3. LLM-as-a-judge – a model grades against a rubric, for qualities like helpfulness or faithfulness to sources.
python
def score_support_reply(case, output):
    checks = {
        "valid_json": is_json(output),
        "cites_order_id": case["order_id"] in output,
        "under_120_words": len(output.split()) <= 120,
    }
    checks["polite_and_correct"] = judge(rubric=RUBRIC, answer=output, context=case["context"]) >= 4
    return checks

Keep judge rubrics specific ("mentions the 30-day refund window") rather than vague ("is good").

Run evals in CI

Treat the eval suite like a test suite:

  • Run it on every change to prompts, models, tools or retrieval.
  • Track scores over time and fail the build on meaningful drops.
  • Store outputs so you can diff behaviour between versions.

Evaluate in production too

Offline evals never cover everything. In production:

  • Run cheap heuristic checks on all traffic.
  • Run expensive judge-based scoring on a sample, such as 5–10% of requests.
  • Collect user feedback signals (thumbs, edits, retries).

Close the loop

The most valuable habit: when a production trace fails a check or a user reports a bad answer, turn it into a new eval case. Your suite then grows from real behaviour instead of guesses.

Agent-specific checks

For agents, score the path as well as the answer:

  • Did it call the right tools, in a sensible order?
  • How many steps and tokens did it take?
  • Did it stop when it should, or loop?

Key takeaways

  • Build a small golden dataset from real usage.
  • Prefer deterministic checks; use LLM judges with specific rubrics.
  • Run evals in CI and sample them in production.
  • Turn every production failure into a new eval case.