The model
A retry repeats work after the caller cannot determine whether an earlier attempt completed. Timeouts create this ambiguity: the server may have committed the operation while the response was lost. Safe retry design begins by defining idempotency, deadlines, and which errors may be transient.
A concrete walk-through
An idempotency key lets the service associate repeated attempts with one logical operation and return the recorded result. Exponential backoff with jitter spreads retries over time, while a retry budget limits amplification. Backpressure slows or rejects new work when queues or downstream services approach capacity.
Costs and failure cases
Retrying every failure can turn a brief outage into a traffic surge. A timeout is not proof of failure, and retries at multiple layers multiply attempts. Queue length without a bound merely converts overload into rising latency and memory use.
Check your understanding
A client retries three times, an API gateway retries twice, and a worker retries four times. Calculate the maximum downstream attempts for one request and propose where to enforce a shared retry budget.
Further reading
AWS Builders’ Library: Making Retries Safe with Idempotent APIs