CUDA GEMM: measure before optimizing

Compare naive, shared-memory tiled, and cuBLAS FP32 multiplication with explicit correctness checks.

The problem

GPU matrix multiplication combines arithmetic, memory reuse, launch overhead, and numerical choices. A faster-looking result is useful only when it computes the intended answer under comparable conditions.

What I built

I implemented naive, tiled-16, tiled-32, and cuBLAS FP32 comparisons on identical seeded inputs. A double-precision CPU reference checks the library result; every implementation is checked before recording timings.

Measured experiment

Who it is for: Learners and GPU engineers comparing matrix-multiplication implementations.

First task: Read the recorded comparison, then run the checked GEMM executable on your own selected GPU.

What you can produce: Correctness checks plus latency/throughput comparisons for naive GEMM, shared-memory tiles, and strict FP32 cuBLAS.

Current scope: Implemented CUDA comparison with recorded RTX A5000 measurements. Speedups depend on shape and hardware; later optimizations remain future work.

Start with the example Source and documentation Step-by-step tutorials

A first useful result

Try a non-tile-aligned 17 × 31 × 23 case, then square 128 and 256 cases. The tutorial offers both a fresh CUDA run and a route that regenerates recorded RTX A5000 plots without a GPU.

Evidence and limits

The recorded comparison specifies tolerances, warmup, samples, tool versions, and provenance. Timings exclude allocation and host transfers; cuBLAS uses FP32 pedantic math with TF32 disabled. Results depend on shape and hardware and are not an end-to-end speedup promise.

A recorded comparison

FP32 GEMM throughput for four implementations at 128 and 256 square shapes

Measured on an RTX A5000 on 2026-10-06, with TF32 disabled. Values use median CUDA-event kernel latency, excluding allocation and transfers. All implementations passed the recorded correctness checks. Raw data and measurement metadata accompany this chart; these small shapes illustrate why tiling does not guarantee a speedup.