Create a benchmark that sums the same matrix with two loop orders. Keep the compiler flags, data, output, warm-up, and environment consistent. Repeat enough times to see run-to-run variation and report the median plus spread rather than one best result.
Use a profiler or hardware counters when available to test whether cache misses changed. If they did not, revise the explanation. Then try tiling and compare the added code complexity with the measured improvement.
Your report should state CPU model, compiler, optimization level, dimensions, measurement method, and limitations. The linked CS:APP Cache Lab offers a deeper simulator-based exercise.