The model
Latency measures how long one operation takes; throughput measures how many operations complete per unit time. A system can have high throughput and poor tail latency if requests wait in queues. Concurrency increases work in flight, but beyond capacity it often increases waiting rather than useful output.
A concrete walk-through
If a service can process 500 requests per second but receives 700, backlog grows until arrivals fall, requests are shed, or the service fails. Little’s Law relates average items in a stable system to arrival rate and time in the system. This gives a consistency check for queue and latency measurements.
Costs and failure cases
Average latency hides p95 and p99 delays that users experience during bursts or dependency stalls. A queue with no upper bound converts overload into memory pressure and long waits. Admission control and load shedding can preserve critical operations when capacity is exceeded.
Check your understanding
A queue contains 2,000 jobs and drains at 250 jobs per second while new work pauses. Estimate the drain time, then explain what changes if arrivals continue at 200 per second.