Your SSDs Are the New DRAM: Co-Activation-Aware KV Cache Offloading Across Multiple Drives
The Memory Wall Hitting LLM Serving
Running large language models at scale is increasingly a memory problem, not a compute problem. As context windows stretch to hundreds of thousands of tokens and multi-user serving demands persistent KV caches across sessions, the bottleneck has shifted decisively: you can have the fastest GPU in the world, but if you can’t feed it KV data fast enough, throughput collapses.
The standard mitigation — offloading KV cache from GPU high-bandwidth memory (HBM) to CPU DRAM — buys time but doesn’t solve the underlying economics. DRAM is expensive, capacity-limited (a server might top out at a few terabytes), and provisioning it purely to buffer KV cache is hard to justify at scale. SSDs offer 10-100x more capacity per dollar, but the conventional wisdom has been that they’re too slow: a naive page-in/page-out scheme over a single NVMe drive quickly becomes bottlenecked on PCIe bandwidth, typically around 7 GB/s for a high-end consumer drive.
Swarm, described in this paper, attacks this problem by asking a different question: what if you stop treating SSDs as a single slow device and start treating a cluster of them as a parallel disaggregated memory tier?
Why Naive SSD Offloading Fails
To understand the contribution, it’s worth unpacking why a straightforward SSD paging scheme underperforms. KV cache access during autoregressive decoding isn’t random — every forward pass needs the keys and values for all prior tokens in the context. For a long-context request, that’s potentially gigabytes of data that must transit from storage to GPU memory before the next token can be generated. With a single NVMe drive at ~7 GB/s sequential read and PCIe 4.0 x4 as the ceiling, you’re looking at hundreds of milliseconds of latency per decode step for large contexts. That turns a fast model into a slow one.
Striping naively across multiple drives (RAID-0 style) would multiply aggregate bandwidth, but it treats all KV blocks as equivalent. Real inference workloads have structure: different attention heads, layers, and token positions have correlated access patterns. Ignoring this means you fetch blocks you don’t immediately need while stalling on ones you do.
Co-Activation as a First-Class Signal
The central insight of Swarm is co-activation awareness — the observation that KV cache blocks are not accessed independently. Certain blocks tend to be retrieved together within a single forward pass, because they correspond to tokens or layers that are jointly attended to. This forms a co-activation graph: edges between blocks that frequently appear in the same fetch batch.
Swarm exploits this graph in two ways. First, it uses co-activation patterns to place blocks across SSDs. Blocks that are frequently co-activated get assigned to different physical devices, so a single forward pass can issue parallel reads across multiple SSDs simultaneously rather than serializing them through one device. Second, the co-activation signal drives prefetching decisions — the system can speculatively load blocks it predicts will be needed in upcoming decode steps based on which blocks are currently hot.
The result is that aggregate bandwidth scales roughly linearly with the number of SSDs, while access latency for correlated block groups is amortized across parallel I/O operations rather than paid sequentially.
Engineering the Multi-SSD Tier
Building a practical system around this idea requires solving several non-trivial engineering problems. The co-activation graph itself must be learned or approximated efficiently — you can’t afford to profile every request exhaustively. Swarm’s approach leans on the structural regularity of transformer attention: layer-wise access patterns are largely model-determined, and across-request patterns can be approximated from prefix sharing and request clustering.
Placement decisions also need to remain coherent under eviction pressure. As new KV cache blocks are written to SSD (when DRAM fills up), the placement algorithm must maintain the invariant that co-activated blocks stay on separate devices, or performance degrades back toward the single-device case. This requires a placement-aware eviction policy rather than simple LRU.
There’s also the question of the GPU-to-SSD pipeline itself. Efficient DMA transfers, asynchronous I/O, and careful overlap between compute and prefetch are necessary to prevent the GPU from stalling while SSD reads complete. Swarm integrates with the inference engine’s scheduling loop to ensure prefetch windows align with decode batch timing.
Implications for Production LLM Infrastructure
The broader significance here is cost. Today, a production serving stack for long-context or persistent-session workloads has to provision substantial DRAM headroom — often many hundreds of gigabytes — purely as a KV buffer. At cloud DRAM pricing, this is a meaningful fraction of total serving cost. An efficient SSD tier that delivers, say, 4x the capacity at a fraction of the price changes the economics of context length substantially.
The multi-SSD approach also has a natural scaling path: adding NVMe drives is modular and incremental in a way that adding DRAM is not. A server with 8 NVMe bays and a smart placement layer can deliver aggregate KV read bandwidth competitive with DRAM while storing orders of magnitude more context.
Watch for follow-on work in a few directions: integrating co-activation-aware placement with speculative decoding (where prefetch prediction becomes even more valuable), extending the model to hierarchical storage tiers (DRAM + SSD + slower storage), and hardware-software co-design as CXL-attached memory and computational storage devices mature. The memory hierarchy problem for LLM inference is far from solved, and Swarm represents a meaningful step toward treating heterogeneous storage as a first-class inference resource rather than an afterthought.