[Ramp-up] Does the warp_pipeline_stage priority actually matter for the 8-wave kernels?
This issue proposes an experiment to verify the impact of warp pipeline stage priorities on the performance of 8-wave GEMM kernels. The goal is to measure TFLOPS and MFMA efficiency under different priority configurations to determine if the current 'memory-outranks-compute' setting is beneficial or inert.