Traditional tests assert exact outputs. AI features rarely produce the same output twice, so teams either skip testing or rely on "it looked fine when I tried it". Both lead to silent regressions when you change a prompt, a model or a retrieval setting.
Evals are the answer: repeatable checks that score AI behaviour on a fixed set of cases.
Start with a golden dataset
Collect 30–100 real examples of the task, each with the input and what a good answer must contain. Pull them from:
- Real user requests (anonymised).
- Known failure cases and bug reports.
- Edge cases: empty input, very long input, other languages, adversarial input.
Small and representative beats large and synthetic.
Choose the right scorer for each check
Use the cheapest scorer that works.
- Deterministic assertions – valid JSON, required fields present, length limits, no banned phrases, correct tool called.
- Reference comparison – does the answer contain the key facts from the expected answer?
- LLM-as-a-judge – a model grades against a rubric, for qualities like helpfulness or faithfulness to sources.
def score_support_reply(case, output):
checks = {
"valid_json": is_json(output),
"cites_order_id": case["order_id"] in output,
"under_120_words": len(output.split()) <= 120,
}
checks["polite_and_correct"] = judge(rubric=RUBRIC, answer=output, context=case["context"]) >= 4
return checksKeep judge rubrics specific ("mentions the 30-day refund window") rather than vague ("is good").
Run evals in CI
Treat the eval suite like a test suite:
- Run it on every change to prompts, models, tools or retrieval.
- Track scores over time and fail the build on meaningful drops.
- Store outputs so you can diff behaviour between versions.
Evaluate in production too
Offline evals never cover everything. In production:
- Run cheap heuristic checks on all traffic.
- Run expensive judge-based scoring on a sample, such as 5–10% of requests.
- Collect user feedback signals (thumbs, edits, retries).
Close the loop
The most valuable habit: when a production trace fails a check or a user reports a bad answer, turn it into a new eval case. Your suite then grows from real behaviour instead of guesses.
Agent-specific checks
For agents, score the path as well as the answer:
- Did it call the right tools, in a sensible order?
- How many steps and tokens did it take?
- Did it stop when it should, or loop?
Key takeaways
- Build a small golden dataset from real usage.
- Prefer deterministic checks; use LLM judges with specific rubrics.
- Run evals in CI and sample them in production.
- Turn every production failure into a new eval case.