← Back to dispatches

Are Your GPU Kernels Actually Fast? Benchmarking Against the Laws of Physics

gpuperformance-engineeringcudainference-optimization

Working from the abstract and paper metadata provided, here is the explainer:


The Problem With Beating a Bad Baseline

Most GPU kernel optimization benchmarks measure the wrong thing. If your baseline is a naive PyTorch implementation and your optimized kernel is 3x faster, that sounds impressive — but what if the hardware could theoretically do 10x better? You’ve beaten software, not physics. For agentic AI systems now capable of generating and rewriting CUDA kernels autonomously, this distinction matters enormously: a benchmark that rewards relative speedup over a flawed reference will happily declare victory on code that still leaves most of the GPU idle.

SOL-ExecBench proposes a different standard. Instead of measuring speedup over a software baseline, it measures proximity to the speed of light — the theoretical maximum throughput a given piece of hardware can deliver for a given computation. This reframes GPU optimization as a physics problem with an absolute ceiling, not a relative improvement contest.

What “Speed of Light” Means in Practice

The speed-of-light (SOL) metric is familiar to anyone who has profiled with NSight Compute. For a given kernel, you compute what the roofline model says is physically achievable given the operation’s arithmetic intensity and the GPU’s memory bandwidth and FLOP ceilings. An SOL score of 85% means the kernel is extracting 85% of what the silicon can theoretically deliver; 40% means there’s significant headroom left.

This is a hard constraint grounded in hardware specs — memory bandwidth, compute throughput, occupancy limits — rather than the accidental performance of whatever reference implementation someone happened to write. It also makes scores comparable across kernels with very different computational profiles: a memory-bound attention kernel and a compute-bound matrix multiplication both get scored against their own relevant ceiling.

The Benchmark Itself

SOL-ExecBench consists of 235 CUDA kernel optimization problems drawn from 124 production and emerging AI models. The coverage is deliberately broad: language models, diffusion models, vision encoders, audio processing, video pipelines, and hybrid architectures. This is not a benchmark of synthetic microkernels — these are the actual workloads running in production or on the frontier of deployment.

The benchmark targets NVIDIA Blackwell GPUs, which is a meaningful choice. Blackwell introduces new architectural features (including FP4 support and updated Tensor Core configurations) that create fresh optimization surface area. Kernels optimized for Ampere or Hopper may behave quite differently on Blackwell, so benchmarking against hardware limits on the current generation keeps the results from aging immediately.

The “ExecBench” framing is also deliberate: this is an execution benchmark. It runs kernels on real hardware and measures real runtimes against real theoretical limits. There’s no proxy metric, no synthetic workload stand-in.

Why This Matters for Agentic AI Systems

The timing of this paper is not coincidental. AI coding agents are now capable of generating non-trivial CUDA kernels and iterating on them through search or reinforcement learning. Systems like these need reward signals, and the reward signal shapes what gets learned.

If you train an agent to maximize speedup over a PyTorch baseline, you get an agent that exploits the inefficiencies of PyTorch — which is valuable, but bounded and potentially misleading. The agent may converge on a kernel that is 5x faster than PyTorch but still only at 30% SOL efficiency. Worse, the agent has no signal that it’s leaving performance on the table.

Benchmarking against hardware limits gives an agent (or a human engineer) a signal that doesn’t saturate prematurely. An SOL score of 70% is a clear indication that there’s 30% remaining — and that 30% has a physical interpretation. The agent knows what it’s chasing.

This also affects how benchmark progress is interpreted externally. A leaderboard sorted by speedup-over-baseline can be gamed by choosing a weak baseline. A leaderboard sorted by SOL efficiency is harder to manipulate — the ceiling is set by NVIDIA’s datasheets, not by whoever wrote the reference kernel.

Practical Implications for Kernel Developers

For engineers working on GPU performance directly, SOL-ExecBench offers a standardized way to audit optimization work against hardware limits rather than against historical implementations. A kernel might pass code review and show strong benchmark numbers while sitting at 45% SOL efficiency — a situation that’s invisible without this kind of analysis.

The 235-problem corpus also provides a structured catalog of real-world kernel shapes across modern AI workloads. Even without the benchmarking framework, this is useful signal about what the actual distribution of production kernels looks like in 2025-2026.

What to Watch

The immediate question is how agentic systems perform against SOL-ExecBench scores versus traditional speedup metrics — and whether the two rankings diverge significantly. If they do, it’s evidence that existing agent-generated kernels are winning on the wrong objective. The secondary question is whether SOL-based reward signals produce measurably better-optimized kernels when used to train or guide optimization agents. That result would validate the benchmark’s design premise and likely push the field toward hardware-efficiency as the primary optimization target.

Generated by claude-sonnet-4-6