This trace follows the actual state transitions behind the companion Latency, Throughput, and Queueing Are Linked. It describes a common execution path; implementation details can vary, so keep the contract separate from the mechanism.
Step 1: Measure arrival rate
Latency measures how long one operation takes; throughput measures how many operations complete per unit time. A system can have high throughput and poor tail latency if requests wait in queues. Concurrency increases work in flight, but beyond capacity it often increases waiting rather than useful output.
Step 2: Track work in flight
If a service can process 500 requests per second but receives 700, backlog grows until arrivals fall, requests are shed, or the service fails. Little’s Law relates average items in a stable system to arrival rate and time in the system. This gives a consistency check for queue and latency measurements.
Step 3: Observe queue growth
When arrivals exceed service capacity, queued work grows and tail latency rises; bound the queue and protect critical operations with admission control.
At this point, record the state that changed and check the invariant before advancing. If the operation repeats, make clear which values persist and which are recomputed.
Step 4: Read tail latency
Average latency hides p95 and p99 delays that users experience during bursts or dependency stalls. A queue with no upper bound converts overload into memory pressure and long waits. Admission control and load shedding can preserve critical operations when capacity is exceeded.
Step 5: Reject or shed excess work
A queue contains 2,000 jobs and drains at 250 jobs per second while new work pauses. Estimate the drain time, then explain what changes if arrivals continue at 200 per second.
The trace is complete when the result satisfies the stated contract. Compare this model with the concrete runtime or system you are studying before making a performance claim.