← Back to dispatches

Cheating Double Precision: How INT8 Emulation Unlocks GPU Speedups for HPC Workloads

performance-engineeringgpuhpcnumerical-computing

I’ll work from the abstract and my knowledge of these topics.


The Problem: FP64 Workloads on INT8-Optimized Hardware

Modern GPUs are not built equally across precisions. The hardware generations that dominate today’s HPC clusters — from NVIDIA’s Volta onward — dedicate dramatically more silicon to low-precision tensor cores (INT8, FP16, BF16) than to full double-precision (FP64) units. An A100, for example, delivers roughly 312 TFLOPS of INT8 throughput versus around 19.5 TFLOPS of FP64 — a 16× gap that grows with each GPU generation.

For graphics, machine learning training, and inference, this is a solved problem: those workloads naturally tolerate or even benefit from lower precision. But for scientific computing — quantum chemistry, materials simulation, fluid dynamics — FP64 is often a hard requirement. The physics breaks if you round too aggressively. That leaves a huge class of HPC workloads effectively stranded on expensive hardware they can barely use at full capacity.

This paper attacks that mismatch directly, asking: can we run FP64 simulations on INT8 hardware without sacrificing correctness, and without forcing scientists to rewrite their code?

The Approach: Precision Emulation via SCILIB-Accel

The authors’ answer is precision emulation — representing double-precision numbers using combinations of lower-precision operations, then offloading those emulated computations to the INT8 tensor cores that dominate modern GPU die area.

The vehicle for this is SCILIB-Accel, an automatic BLAS offload layer that intercepts standard BLAS calls (think DGEMM for matrix multiplication) and transparently reroutes them to GPU-accelerated emulated equivalents. The key word is transparent: the target application, the LSMS code in the MuST (Multiple Scattering Theory) suite for ab initio electronic structure calculations, requires zero source changes. Scientists running complex quantum mechanical simulations don’t need to touch their Fortran.

The memory model underpinning this is cache-coherent Unified Memory Architecture — the CPU and GPU share a common address space, so SCILIB-Accel can intercept matrix data in-place without explicit host-to-device copies scattered through the application code. This matters because LSMS wasn’t written with GPU offloading in mind; unified memory is what makes the “no code changes” promise credible.

The Precision Problem Is Operator-Dependent

Here’s where it gets interesting from an engineering standpoint: the paper finds that emulation accuracy isn’t just a function of the precision scheme — it also depends on the properties of the matrices being multiplied.

This is a crucial insight. In electronic structure calculations, you encounter a range of operators. Some are well-conditioned — their matrix entries span a relatively narrow range and behave nicely under rounding. Others are ill-conditioned, with large condition numbers where small floating-point errors in intermediate products cascade into significant errors in the result. A fixed INT8 emulation scheme that works fine for one operator may silently degrade accuracy for another.

The authors address this with tunable precision emulation: the emulation scheme can be adjusted per-operator (or per-class-of-operators) to dial in the right trade-off between speed and accuracy. This is reminiscent of mixed-precision strategies in ML training, where practitioners selectively keep certain layers in FP32 while running others in FP16 — except here the tuning is driven by mathematical conditioning rather than gradient stability.

Why This Architecture Matters for Developers

If you build or maintain HPC software, the architectural pattern here deserves attention regardless of your domain.

The BLAS interception layer approach is a proven strategy — it’s the same principle behind Intel MKL’s transparent replacement of reference BLAS, or drop-in cuBLAS wrappers. What SCILIB-Accel adds is the precision emulation logic sitting between the BLAS interface and the actual GPU kernels. This creates a clean separation: application code calls standard DGEMM, the interception layer decides what precision strategy to apply, and the GPU kernel executes accordingly. For legacy scientific codebases where rewriting is politically or practically impossible, this is a compelling path to GPU utilization.

The unified memory angle is also worth watching. As architectures like AMD’s MI300X and future Grace-Hopper successors push toward truly coherent CPU-GPU memory pools, the overhead that historically made transparent offloading expensive (explicit synchronization, redundant copies) decreases. SCILIB-Accel’s design is well-positioned for this trend.

Implications

The broader takeaway is that the precision gap between what modern GPUs are optimized for and what scientific workloads require doesn’t have to be bridged by rewriting applications. Software-layer emulation, applied through standard library interfaces, can extract meaningful acceleration from tensor core hardware while preserving numerical integrity — as long as the emulation is tunable and informed by the mathematical properties of the computation.

For developers working on HPC middleware, numerical libraries, or scientific application frameworks, this paper is a concrete proof-of-concept that the BLAS layer remains a powerful leverage point. The combination of automatic offloading, coherent memory, and precision-aware emulation may become a standard pattern as the hardware gap between INT8 and FP64 continues to widen.

The full paper is available on arXiv.

Generated by claude-sonnet-4-6