← Back to dispatches

Teaching AI to Write CUDA Kernels That Beat the Compiler

cudaperformance-engineeringinference-optimization

Why LLMs Keep Losing to Compilers at CUDA

If you’ve ever tried to squeeze real performance out of a GPU, you know the gap between “runs on GPU” and “runs fast on GPU” is enormous. Writing CUDA kernels that beat PyTorch’s built-in ops requires intimate knowledge of memory hierarchy, warp scheduling, shared memory bank conflicts, and a dozen other hardware-level concerns that take years to internalize. torch.compile and Triton have made this more accessible, but the ceiling for hand-tuned CUDA remains significantly higher than what compiler-based tools consistently achieve.

The natural question is whether LLMs can close that gap. The frustrating answer, until recently, has been no—and not for lack of trying.

The Two Dead Ends

Existing approaches to LLM-based CUDA generation fall into two camps, and both hit a wall.

The first is training-free refinement: prompt an LLM, run the kernel, feed back the profiling output or error message, and iterate. This works for simple kernels but struggles to escape local optima. Without updating the model’s weights, you’re relying on in-context reasoning to navigate an optimization landscape that experts spend careers learning. The model doesn’t learn from the thousands of failed attempts—each session starts from scratch.

The second camp fine-tunes models within fixed multi-turn execution-feedback loops. You collect (problem, solution, feedback) trajectories and train on them. This is better, but the “fixed loop” structure is a critical limitation. The model learns to follow a rigid script: generate, observe error, patch, repeat. It doesn’t develop genuine search strategies or learn when to abandon an approach and rethink from first principles. The exploration is shallow by design.

Both paradigms, as the paper puts it, “fail to fundamentally improve the model’s optimization capabilities.”

Agentic RL as an Escape Hatch

CUDA Agent takes a different route: large-scale agentic reinforcement learning. Rather than defining a fixed interaction loop, the system treats CUDA optimization as an open-ended agent task. The model can issue arbitrary sequences of actions—writing code, running profilers, reading hardware documentation, comparing kernel variants, adjusting tile sizes—and receives reward signals based on actual kernel speedup over a reference implementation.

The key architectural shift is that the agent is trained to search, not just refine. There’s no predetermined number of turns, no fixed schema for how feedback gets incorporated. The model learns, through millions of rollouts, which investigative strategies actually pay off on real hardware.

This is a meaningful departure from supervised fine-tuning on expert trajectories, where the model imitates what a good engineer does. RL on execution feedback lets the model discover strategies that human experts might not articulate or even consciously use.

What “Large-Scale” Actually Means Here

The “large-scale” in the title isn’t decorative. Training agentic RL for code optimization is expensive in a specific way: every rollout requires compiling and executing CUDA code on real GPU hardware, not just running a cheap simulator. Reward computation is slow and noisy. Kernels that compile correctly may produce wrong outputs (numerical issues are common with aggressive optimizations like float16 accumulation or reordered memory accesses), and distinguishing correctness failures from performance failures requires careful infrastructure.

The paper’s training setup had to solve distributed kernel execution at scale—running thousands of candidate kernels in parallel across GPU clusters while maintaining reproducible timing measurements, which are notoriously sensitive to thermal state, competing processes, and driver behavior.

The Benchmark Gap

The paper benchmarks against torch.compile as the primary baseline, which is the right comparison. torch.compile with the Inductor backend is not a pushover—it applies fusion, tiling, and vectorization passes that typically beat naive PyTorch by 2–5x on common operators.

CUDA Agent trained models show competitive or superior performance across a range of kernels including matrix multiplications, attention variants, and element-wise fused operations. The gap is most pronounced on non-standard kernels that fall outside the well-optimized paths that compilers know about—exactly the regime where human experts writing custom CUDA earn their keep, and where prior LLM approaches struggled most.

What This Means for Developers

A few practical implications worth tracking:

Custom op development gets cheaper. Right now, writing a high-performance custom CUDA kernel for a novel operation (a new attention variant, a fused normalization scheme, an unusual reduction) requires either deep expertise or accepting significant performance overhead. Agentic RL systems that can search the optimization space effectively could compress this from a weeks-long expert task to a tool-assisted iteration cycle.

The feedback loop is the asset. The paper’s core insight applies broadly: if you’re building LLM-based code generation for performance-sensitive tasks, the quality of your execution environment and reward signal matters more than model scale. A smaller model trained with real hardware feedback will outperform a larger model trained on synthetic data or fixed interaction patterns.

Triton may be the better integration point. Triton’s higher abstraction level makes it easier to generate correct-by-construction kernels that the compiler can optimize, potentially making it a more tractable target for LLM generation than raw CUDA. Watch for follow-up work that applies agentic RL at the Triton level, where the search space is smaller and correctness checking is simpler.

The deeper trend here is that LLMs reaching expert-level performance on specialized technical tasks requires training regimes that let models accumulate real experience in the task environment—not just imitate experts from the outside.

Generated by claude-sonnet-4-6