This trace follows the actual state transitions behind the companion Profile Before Rewriting the Slow Path. It describes a common execution path; implementation details can vary, so keep the contract separate from the mechanism.
Step 1: State a bottleneck hypothesis
Profiling attributes resource use to parts of a running program; benchmarking compares defined workloads under controlled conditions. A useful optimization starts with a hypothesis about the bottleneck and gathers evidence at the level where the cost occurs.
Step 2: Run a representative workload
If a CPU profile shows most samples inside JSON parsing, tuning database indexes will not fix the measured CPU bottleneck. A benchmark should use representative data, warm-up behavior, stable machine conditions, and enough repetitions to show variation rather than one favorable run.
Step 3: Collect profile or timings
A profile showing time in parsing is evidence for a CPU bottleneck; measure the real request again because a local optimization may not affect total latency.
At this point, record the state that changed and check the invariant before advancing. If the operation repeats, make clear which values persist and which are recomputed.
Step 4: Change one dominant cost
Instrumentation adds overhead and can distort timing, so use the least intrusive tool that answers the question. Microbenchmarks isolate operations but may fail to predict full-system behavior. Preserve correctness checks and compare end-to-end latency after any local speedup.
Step 5: Re-measure end-to-end behavior
A new implementation is twice as fast in a microbenchmark but the endpoint latency is unchanged. Name three possible reasons and the next measurement you would collect.
The trace is complete when the result satisfies the stated contract. Compare this model with the concrete runtime or system you are studying before making a performance claim.