Interview prompt
Explain retries need an operation contract to an engineer who understands the surrounding system but has not used this technique. Walk from its contract to a concrete operation, then discuss where it fails or becomes expensive.
A strong answer
A retry repeats work after the caller cannot determine whether an earlier attempt completed. Timeouts create this ambiguity: the server may have committed the operation while the response was lost. Safe retry design begins by defining idempotency, deadlines, and which errors may be transient.
An idempotency key lets the service associate repeated attempts with one logical operation and return the recorded result. Exponential backoff with jitter spreads retries over time, while a retry budget limits amplification. Backpressure slows or rejects new work when queues or downstream services approach capacity.
A complete answer also calls out the assumptions that control correctness. Retrying every failure can turn a brief outage into a traffic surge. A timeout is not proof of failure, and retries at multiple layers multiply attempts. Queue length without a bound merely converts overload into rising latency and memory use.
Close by describing one representative test or measurement. A client retries three times, an API gateway retries twice, and a worker retries four times. Calculate the maximum downstream attempts for one request and propose where to enforce a shared retry budget.
Follow-up questions
Answer the follow-ups in the frontmatter. Use the linked article for the concept and the trace to make the explanation concrete.