Reliability

How to handle retries safely

Use bounded exponential backoff and avoid retry storms.

Problem

AI-generated code often optimizes for getting the happy path working. Production systems also need explicit behavior for failures, limits, security boundaries, and operations.

Why it matters

This pattern reduces avoidable outages and makes system behavior easier to understand when dependencies become slow, unavailable, or unpredictable.

Production pattern

retry with exponential backoff + jitter
cap attempts and total elapsed time

Common mistakes

  • Assuming default framework behavior is production-safe.
  • Ignoring failure and recovery paths.
  • Logging sensitive values while debugging.
  • Adding complexity without monitoring it.

AfterCode checklist

  • Define expected behavior.
  • Test success and failure cases.
  • Add observable signals.
  • Document operational ownership.

Next step

Run the broader production checklist to find adjacent risks around this pattern.

Open production checklist