What It Actually Takes to Run the World's Fastest Supercomputer at Full Speed
I wasn’t granted permission to fetch the paper. I’ll write the explainer from the abstract and my knowledge of Aurora’s public benchmark results.
Getting an exascale supercomputer to actually run at exascale — and keep doing it reliably — turns out to be a much harder engineering problem than getting it to do so once for a press release. The paper “Sustaining Exascale Performance: Lessons from HPL and HPL-MxP on Aurora” documents exactly that gap: the difference between peak capability and sustained production performance, as learned through three successive benchmark campaigns on Aurora at Argonne National Laboratory.
What Aurora Is, and Why It’s a Hard Target
Aurora is unusual even among exascale machines. It’s the first large-scale deployment of Intel discrete GPUs — the Intel Data Center GPU Max 1550 (Ponte Vecchio) — across 10,624 compute nodes, each pairing two Intel Xeon CPU Max processors with six of those GPUs. The interconnect is HPE Cray’s Slingshot-11, and notably the NICs are CPU-attached rather than GPU-attached. That last detail matters: moving data between GPU memory and the network requires an extra hop through host memory, which becomes a latency and bandwidth bottleneck at scale.
The system also holds the record as the largest production Slingshot-11 deployment. Dragonfly topology at that scale means routing decisions, congestion management, and link reliability all interact in ways that only manifest when you’re using essentially the entire machine.
The Benchmarks: HPL and HPL-MxP
HPL (High Performance LINPACK) is the benchmark that determines the TOP500 ranking. It solves a dense system of linear equations using LU factorization in 64-bit floating point — a deeply communication-bound workload once you’re at this scale. Every node must participate in panel broadcasts and trailing matrix updates across the full machine. A single slow node, link flap, or memory error can drag down the entire run.
HPL-MxP is the mixed-precision variant. It uses lower-precision arithmetic (typically BF16 or FP16) for the compute-heavy matrix multiplications and higher precision for iterative refinement. This is important because it maps well to the tensor-core-style hardware that modern AI-oriented accelerators (including Intel’s Ponte Vecchio) are optimized for. HPL-MxP scores are typically 3–5× higher than HPL on the same hardware, reflecting the hardware’s actual peak math throughput rather than its FP64 ceiling.
Aurora achieved over 1 exaflop on HPL — making it the first US exascale system — and substantially higher throughput on HPL-MxP. The paper’s contribution isn’t announcing those numbers; it’s explaining what it took to reliably reproduce them.
Where the Real Engineering Lives
The paper’s framing around “three successive campaigns” is telling. The first campaign presumably established that the system could hit exascale. The subsequent ones are about understanding why performance varied and building the operational practices to control it.
Several categories of challenge are canonical to runs at this scale:
Node health and exclusion. With over 10,000 nodes, some will always be in a degraded state — GPUs with correctable error rates climbing toward uncorrectable, NICs with marginal links, DRAM pages going bad. The HPL run has no fault tolerance: a hardware fault mid-run means restarting from scratch. So pre-run health screening, automated node exclusion policies, and the definition of “healthy enough to include” become critical operational decisions. Being too conservative leaves performance on the table; too permissive and you waste hours on a doomed run.
CPU-attached networking and data movement. Because Aurora’s NICs are CPU-attached rather than GPU-attached, HPL’s communication phases require explicit staging through host memory. At exascale, the panel broadcast — the column of the LU factorization distributed to every node — is one of the most latency-sensitive operations. The paper likely details tuning of MPI collectives, buffer pinning strategies, and how the software stack was adjusted to hide this architectural bottleneck.
GPU software maturity. Intel’s oneAPI stack (oneMKL, MPI, Level Zero) was maturing alongside Aurora’s deployment. Running HPL on Intel GPUs at this scale was itself a validation exercise for the entire software ecosystem. Compiler bugs, library regressions, and driver issues that are inconsequential in small tests can manifest as correctness or performance failures at scale.
Interconnect at the limit. Slingshot-11 at this deployment size means tens of thousands of links, adaptive routing under heavy all-to-all traffic, and congestion that can cascade unpredictably. Tuning traffic patterns, collective algorithms, and even job placement across the dragonfly topology are all levers that matter.
Coordination Across System Layers
The paper’s framing — “coordination across system layers” — is what separates this from a pure benchmarking report. Sustaining performance requires the firmware, OS, runtime, MPI, math libraries, and application to all be tuned together, and for the operational team to understand the dependency chain. A firmware update that improves one metric can break an assumption baked into the MPI implementation. An OS scheduler change can affect GPU utilization patterns in ways that only show up at scale.
This kind of cross-layer knowledge is hard to document and harder to transfer. Papers like this one serve as institutional memory for the HPC community.
What to Watch For
Aurora’s architecture is a preview of where HPC is heading: heterogeneous systems where AI-optimized accelerators are first-class compute, and where sustaining advertised performance in real workloads requires continuous engineering effort rather than one-time tuning.
The lessons here — health screening at scale, software stack co-evolution, CPU-NIC coupling tradeoffs, and the operational discipline required to sustain exascale — will apply directly to the next generation of systems from any vendor. If you’re building distributed systems at scale, the patterns are recognizable: the gap between theoretical throughput and reliable sustained throughput is always larger than the spec sheet implies, and closing it requires owning the full stack.