← Back to dispatches

The GPU Workload That Rewrites Molecular Physics: Profiling ML Force Fields

performance-engineeringgpuscientific-computingsystemsml-inference

Why GPU Efficiency Is the Gating Factor for Next-Gen Simulation

Physics-based simulation sits at the center of drug discovery, materials design, and protein engineering, but for decades the accuracy-cost tradeoff has been brutal. Density functional theory (DFT) gives you quantum-accurate forces, but scales as O(N³) in system size. Classical force fields scale linearly, but they can’t describe bond breaking, charge transfer, or reactions outside their parameterization set. Machine learning force fields (MLFFs) promise to close that gap — near-quantum accuracy at a cost orders of magnitude below DFT. The question that this paper actually answers is: how well do today’s GPUs handle that promise in practice?

The answer turns out to be “not as well as you might hope, and for reasons that reveal deep mismatches between MLFF computation and GPU architecture.”

What Makes MLFF Workloads Structurally Different

Classical force fields (CFFs) like AMBER or CHARMM were essentially co-designed with GPU architectures. Their compute kernels map cleanly to the GPU programming model: neighbor lists enable regular memory access patterns, bonded interactions are embarrassingly parallel, and long-range electrostatics via particle-mesh Ewald decomposes into FFTs that GPUs handle well. The result is high arithmetic intensity and excellent hardware utilization.

MLFFs break almost every one of these favorable properties. Modern architectures — think NequIP, MACE, or SchNet — represent atomic systems as graphs where nodes are atoms and edges encode neighbor relationships. Force inference runs as iterated message passing over this graph: each atom aggregates information from its neighbors, transforms it through a neural network, and passes updated representations outward. The computation is inherently irregular.

Three specific structural problems dominate:

Variable-degree graphs. Atoms in a protein active site may have 5 neighbors; atoms in a bulk solvent region may have 30. GPU threads executing in lockstep (warps of 32) stall on the slowest thread in the group, so high-degree variance translates directly to wasted cycles. Unlike convolution over a dense image, you can’t pad your way to regularity without introducing phantom interactions.

Scatter-reduce operations. Aggregating neighbor messages back to a central atom requires atomic additions to shared memory or global memory — classic bottlenecks that serialize what should be parallel. At large system sizes, contention around popular atoms (e.g., a central metal ion in a binding pocket) becomes measurable.

Equivariant feature transforms. State-of-the-art MLFFs enforce SE(3) or E(3) equivariance — the prediction must rotate when you rotate the molecule. This is physically correct and dramatically improves data efficiency during training, but it requires operations over spherical harmonics and Clebsch-Gordan tensor products that have poor cache locality and don’t map naturally to the matrix-multiply units (Tensor Cores) that dominate modern GPU performance.

The Compute vs. Memory Bottleneck Shift

For CFFs, the dominant cost is typically floating-point arithmetic — you’re doing a lot of per-pair potential evaluations. MLFFs invert this: the bottleneck shifts toward memory bandwidth. Each message-passing step must load neighbor features, which for equivariant models can be high-dimensional tensors, from memory for every edge in the graph. At inference time (unlike training), batch sizes are typically one simulation frame, so you can’t amortize memory latency across a large batch the way you would in, say, transformer inference.

This memory-bound regime is significant for hardware selection. The GPU characteristics that matter most for MLFF workloads — HBM bandwidth, L2 cache size, memory latency — are different from the tensor-compute throughput (FLOPS) metrics that typically dominate GPU marketing and benchmarking for deep learning.

The Software Stack Gap

The paper frames MLFFs as “emerging workloads,” and the framing is apt: the software ecosystem hasn’t caught up. Libraries like OpenMM and GROMACS have decades of GPU kernel optimization for classical simulations. MLFF inference is often dispatched through PyTorch or JAX, which are general-purpose ML frameworks not tuned for the small-graph, high-equivariance regime these models operate in.

Practically, this means a significant fraction of GPU time is spent in framework overhead — kernel launch latency, graph capture, Python dispatch — rather than actual simulation computation. At smaller system sizes (hundreds of atoms), this overhead can dominate the entire runtime, making GPUs only marginally faster than CPUs.

What to Watch For

The characterization laid out in this work points toward several active fronts:

Custom CUDA kernels for equivariant operations. Projects like e3nn and cuEquivariance are beginning to provide fused kernels for Clebsch-Gordan products that approach Tensor Core utilization. Expect this to be a key performance lever over the next two to three years.

Graph-aware batching strategies. Grouping simulation frames such that graphs have similar degree distributions — or padding to fixed-degree graphs with masking — is a tractable engineering problem that could recover significant warp efficiency.

Hardware co-design pressure. As MLFF-driven MD scales up in industrial drug discovery pipelines (where it’s already in production use at several major pharma companies), the workload profile documented in papers like this one will inform what GPU vendors prioritize. The irregular sparse-graph compute pattern shared by MLFFs and graph neural network inference more broadly is increasingly being cited in architecture discussions.

The deeper implication is that MLFF adoption is not just a modeling problem — it’s a systems problem. Characterizing these workloads with the same rigor that HPC groups have applied to classical MD is a prerequisite for closing the gap between MLFF’s theoretical promise and its practical throughput at production scale.

Generated by claude-sonnet-4-6