Monitoring a normal service tells you that something went wrong. With AI agents you also need to know why: which step went off course, what the model saw and which tool returned something unexpected.
That is the job of agent observability: capturing the full decision path of an agent run.
What to capture
Model each agent run as a trace made of spans, the same idea as distributed tracing.
- Model calls – model name, prompt (or a reference to it), output, token counts, latency, cost.
- Tool calls – tool name, arguments, result size, errors, duration.
- Retrieval steps – query, documents returned, scores.
- Decisions – plans, retries, hand-offs between agents.
- Outcome – success, failure, user feedback.
Use OpenTelemetry where you can
OpenTelemetry has semantic conventions for generative AI, so traces can flow into your existing tools rather than a separate silo.
from opentelemetry import trace
tracer = trace.get_tracer("support-agent")
with tracer.start_as_current_span("tool.search_orders") as span:
span.set_attribute("tool.args.customer_id", customer_id)
result = search_orders(customer_id)
span.set_attribute("tool.result.count", len(result))The metrics that matter
- Task success rate – from evals, user feedback or explicit completion signals.
- Steps per task and tokens per task – rising numbers often mean the agent is looping or confused.
- Tool error rate by tool – one flaky tool can derail every run.
- Cost per successful task – more useful than cost per request.
- Latency percentiles – agents with many steps have long tails.
Failure patterns traces reveal
- Loops – the same tool called with the same arguments repeatedly.
- Wrong tool – a similar-sounding tool chosen because descriptions overlap.
- Context overflow – a large tool result pushes out the instructions.
- Silent tool failure – a tool returned an empty result and the model invented an answer.
Privacy and retention
Prompts and outputs can contain personal data. Redact sensitive fields before export, set retention limits and restrict who can read raw traces.
Connect observability to evals
Failed production traces are the best source of new eval cases. Make it one click to turn a bad trace into a test.
Key takeaways
- Trace every model call, tool call and retrieval step.
- Prefer OpenTelemetry-compatible instrumentation.
- Watch steps, tokens and cost per successful task.
- Redact sensitive data and feed failures back into evals.