Observability

How to monitor AI application latency

Track model, retrieval and tool-call latency separately.

Problem

AI-generated code often optimizes for getting the happy path working. Production systems also need explicit behavior for failures, limits, security boundaries, and operations.

Why it matters

This pattern reduces avoidable outages and makes system behavior easier to understand when dependencies become slow, unavailable, or unpredictable.

Production pattern

measure retrieval_ms, model_ms, tool_ms, total_ms

Common mistakes

  • Assuming default framework behavior is production-safe.
  • Ignoring failure and recovery paths.
  • Logging sensitive values while debugging.
  • Adding complexity without monitoring it.

AfterCode checklist

  • Define expected behavior.
  • Test success and failure cases.
  • Add observable signals.
  • Document operational ownership.

Next step

Run the broader production checklist to find adjacent risks around this pattern.

Open production checklist