Write two loops that sum every element in the same row-major matrix: one with columns in the inner loop and one with rows in the inner loop. Keep the work and data identical so the access order is the main variable.
Predict which version reuses each fetched cache line more effectively. Then measure both with repeated runs and, if available, cache counters or the linked CS:APP Cache Lab. Record the machine, compiler, optimization flags, matrix size, and measurement method. Do not present one result as a universal speed ratio.
As a follow-up, tile the matrix so a small block stays in cache while both dimensions are traversed. Explain why the best tile size depends on the cache hierarchy and other active data.