Inside the Black Box: Reverse-Engineering NVIDIA's GPU Command Streams
I don’t have WebFetch access yet, so I’ll write the explainer from the abstract and my knowledge of this domain. Here it is:
The Black Box Between Your Code and Your GPU
Every CUDA developer has stared at a nsys profile wondering: why does this kernel launch take 50 microseconds before execution even starts? Why does a simple cudaMemcpy involve this much CPU time? The answers live in code you’re not allowed to read — NVIDIA’s closed-source userspace driver, libcuda.so. A new paper from arxiv cracks that open, recovering the actual hardware command streams the driver emits and making the full CPU-to-GPU translation path visible for the first time in a systematic way.
How GPU Command Submission Actually Works
When you call cudaLaunchKernel, the CUDA Runtime hands off to the CUDA Driver API, which hands off to NVIDIA’s proprietary userspace driver. That driver’s job is to encode your intent as a sequence of hardware commands — method/data word pairs — and push them into a pushbuffer, a ring buffer in pinned memory that the GPU’s channel host engine continuously reads. The GPU doesn’t understand “launch kernel”; it understands a specific sequence of register writes, pointer loads, and control words that collectively set up and trigger execution.
This encoding is entirely the userspace driver’s responsibility. NVIDIA’s open-source kernel module (open-gpu-kernel-modules, released in 2022) handles only the privileged side — channel allocation, memory mapping, interrupt routing — leaving the pushbuffer construction fully opaque. Nouveau, the open-source reverse-engineered driver, has mapped parts of this for older architectures, but coverage is incomplete and lags modern hardware significantly.
What the Paper Does
The paper establishes a methodology for intercepting and decoding these command streams as they flow from the closed-source driver into hardware. Rather than static binary analysis of libcuda.so (which is heavily obfuscated and changes with every driver release), the approach observes the live command stream — essentially eavesdropping on the pushbuffer traffic during real CUDA workloads.
The recovered streams reveal the concrete hardware operations behind familiar CUDA abstractions:
- Kernel launches decompose into a sequence that includes QMD (Queue Meta-Data) structure writes, grid/block dimension encoding, shader program pointer setup, and a channel kickoff command — all before the GPU scheduler sees the work.
- Memory copies (
cudaMemcpy) are not monolithic; the driver selects between several DMA engines and copy methods depending on transfer size, direction, and whether the source/destination is managed memory. A single API call can emit dozens of commands across multiple hardware engines. - Synchronization primitives like
cudaStreamSynchronizetranslate into semaphore acquire/release sequences with explicit fence operations, revealing how the driver coordinates between the compute engine, copy engines, and the host CPU.
Why This Is Hard, and What Makes It Tractable
The core challenge is that pushbuffer command encodings are not publicly documented. NVIDIA publishes some information through the cl*.h class headers included in their kernel module source, but these cover only a fraction of the command space, and the userspace driver uses many undocumented methods.
The paper’s approach leverages the fact that while the driver is closed, the hardware state it produces is observable. By instrumenting the boundary between the userspace driver and the kernel module — the ioctl interface — and correlating with GPU hardware performance counters and memory traces, the authors can reconstruct command intent even without full documentation. They also cross-reference against Nouveau’s reverse-engineering work and the partial class headers to annotate recovered streams.
The result is a labeled corpus of command streams paired with the CUDA API calls that generated them, across representative workloads including matrix multiplication (cuBLAS), convolution (cuDNN), and simple custom kernels.
Practical Implications for Developers
Performance attribution gets more honest. Profilers today show you “kernel execution time” and “memory transfer time,” but attribution for CPU-side driver overhead is coarse. Understanding that a kernel launch requires encoding a 64-byte QMD structure, writing it to pinned memory, and issuing several channel commands before the GPU scheduler is even involved changes how you think about launch overhead — and why batching small kernels matters more than the GPU execution time alone suggests.
Custom runtime development becomes more grounded. Projects like Triton, JAX’s PJRT backend, and various ML compilers that bypass the CUDA Runtime are essentially reimplementing parts of what libcuda.so does. Until now, that reimplementation involved significant guesswork about which command sequences are actually necessary versus conservative. A documented command corpus provides a ground truth to validate against.
Security and correctness research. GPU command streams are a capability boundary — the kernel module trusts that userspace submits valid commands. Understanding what “valid” looks like is prerequisite to understanding what “invalid but accepted” might look like. This work provides a systematic baseline.
Alternative driver development. Mesa’s NVK Vulkan driver and Nouveau have made significant progress on newer NVIDIA architectures since the kernel module open-sourcing, but CUDA-level compatibility remains out of reach without understanding the CUDA-specific command encoding. This kind of systematic stream recovery is the necessary groundwork.
What to Watch
The immediate value is analytical — a researcher’s tool for understanding behavior, not a drop-in driver replacement. But the methodology matters as much as the current results. If command stream recovery can be kept current across driver versions (a significant engineering challenge given NVIDIA’s release cadence), it becomes infrastructure that the open-source GPU software ecosystem has lacked for years. Watch for follow-on work that applies this to specific performance anomalies, and for integration with existing profiling toolchains that could surface this information to developers without requiring them to read raw command streams themselves.