This trace follows the actual state transitions behind the companion Retries Need an Operation Contract. It describes a common execution path; implementation details can vary, so keep the contract separate from the mechanism.
Step 1: Set the operation deadline
A retry repeats work after the caller cannot determine whether an earlier attempt completed. Timeouts create this ambiguity: the server may have committed the operation while the response was lost. Safe retry design begins by defining idempotency, deadlines, and which errors may be transient.
Step 2: Classify the failure outcome
An idempotency key lets the service associate repeated attempts with one logical operation and return the recorded result. Exponential backoff with jitter spreads retries over time, while a retry budget limits amplification. Backpressure slows or rejects new work when queues or downstream services approach capacity.
Step 3: Check deduplication state
A timeout leaves completion uncertain, so check idempotency before retrying; bound attempts and in-flight work to prevent retries from multiplying overload.
At this point, record the state that changed and check the invariant before advancing. If the operation repeats, make clear which values persist and which are recomputed.
Step 4: Back off with jitter
Retrying every failure can turn a brief outage into a traffic surge. A timeout is not proof of failure, and retries at multiple layers multiply attempts. Queue length without a bound merely converts overload into rising latency and memory use.
Step 5: Enforce a retry or queue budget
A client retries three times, an API gateway retries twice, and a worker retries four times. Calculate the maximum downstream attempts for one request and propose where to enforce a shared retry budget.
The trace is complete when the result satisfies the stated contract. Compare this model with the concrete runtime or system you are studying before making a performance claim.