Diagnosing a Slow Service With Evidence
The lesson starts with an API whose tail latency worsens while average database CPU looks normal. A distributed trace reveals a long wait before query execution. A timeline then compares pool-acquisition time, database duration, downstream calls, and serialization.
The narrator shows how to form and test one hypothesis, examine queue depth and connection state, and repeat the measurement under the same workload. The closing scene warns that increasing pool size can move the queue into the database rather than remove it. Companion material: measurement article, pool-wait trace, and measurement flow.
Related articles
Measure the Work Before You Optimize It
Use workload shape, latency distributions, profiling, and controlled experiments to find the cause of slow software.
HTTP/2 Multiplexing and Connection Management
How HTTP/2 fixes head-of-line blocking with binary framing and multiplexing, and why connection reuse matters for latency.
Latency, Throughput, and the Cost of Coordination
Every system design trade-off is ultimately a balance between doing work fast, doing work often, and paying the cost of making multiple components agree.
New lessons by email
Get new articles and notes on the systems behind everyday software.
One technical dispatch per week. No noise.
Not started
Sign in to save your learning progress.