← Back to dispatches

Split to Win: Why Hybrid LLM Inference Needs Disaggregated Prefill and Decode Stacks

inference-optimizationmamballm-servingperformance-engineering

Working from the abstract and my knowledge of this architecture space, here’s the explainer:


The Problem With One Chip Doing Everything

Modern LLM inference has a split personality, and hardware hasn’t caught up. When you send a prompt to a language model, two fundamentally different workloads execute back-to-back: prefill, where the model digests your input tokens in parallel (compute-bound, GPU-friendly), and decode, where it generates one token at a time (memory-bandwidth-bound, bottlenecked by loading weights on every step). Running both on the same GPU has always been a compromise, but hybrid Mamba-Transformer architectures make that compromise significantly worse.

Hybrid models interleave standard attention layers with State Space Model (SSM) layers — specifically Mamba-style recurrences. SSMs are appealing because they avoid the quadratic attention cost and produce a compact recurrent state, making them efficient in theory. In practice, their element-wise operations and recurrent structure fit poorly on matmul-centric accelerators designed around tensor cores. The result: a model that is neither efficiently prefilling nor efficiently decoding on any single piece of hardware.

DUET is a direct answer to this mismatch.

Disaggregation as a First-Class Design Principle

The core idea in DUET is architectural disaggregation — not just separating prefill and decode scheduling (as systems like Splitwise or DistServe do at the software level), but designing distinct accelerator packages for each phase, matched to the actual compute profile of each.

For prefill, workloads are dominated by large matrix multiplications: projecting keys and values, computing attention scores across long sequences, running feed-forward layers. These map cleanly to high-throughput tensor core operations. DUET’s prefill package is built around this: high FLOPS, high compute intensity, with enough HBM bandwidth to feed the matmul engines efficiently.

Decode is a different animal. At each step you’re loading the full model weight matrix to produce a single output vector — utilization of compute units collapses, and the bottleneck shifts entirely to memory bandwidth. But hybrid models add another wrinkle: the SSM recurrence involves element-wise updates to a state vector. These operations have low arithmetic intensity and are effectively useless on tensor cores. DUET’s decode package prioritizes bandwidth and efficient execution of these irregular, element-wise operations rather than raw FLOPS.

Why Hybrid Models Break the Old Compromise

Pure Transformer inference already suffers from the prefill/decode asymmetry, but the community has developed workable mitigations: continuous batching, chunked prefill, speculative decoding. Hybrid Mamba-Transformer models undermine some of these because the SSM layers don’t batch the same way attention does. The recurrent state is per-sequence and must be maintained and updated independently — you can’t amortize it across the batch the way you can with batched matrix multiplications in attention.

This means that on a homogeneous GPU cluster, you’re constantly paying a tax: either your hardware is over-provisioned for the memory-bandwidth requirements of decode, or it’s under-provisioned for the compute requirements of prefill. The SSM element-wise operations add a third dimension of mismatch that doesn’t exist in pure Transformer models.

DUET’s disaggregated design means each phase runs on purpose-built silicon, eliminating this tax. The prefill package never has to handle SSM recurrences efficiently; the decode package never has to pretend it’s a high-FLOPS matrix engine.

Practical Implications for Inference Systems

For developers building on top of LLM inference infrastructure, DUET points toward a few near-term shifts worth tracking:

Hardware heterogeneity in serving clusters will increase. If disaggregated inference becomes the norm for hybrid models, serving stacks will need to treat the prefill and decode stages as separate, schedulable units with different hardware affinities. Frameworks like vLLM and SGLang are already moving toward disaggregated serving abstractions, but hardware-level disaggregation raises the stakes for orchestration logic.

SSM-specific acceleration is a real design constraint. The paper’s framing makes explicit what practitioners have suspected: generic tensor-core accelerators are a poor fit for the recurrent, element-wise compute in SSM layers. This has implications not just for hardware but for model design — hybrid ratios (how many SSM vs. attention layers) will increasingly be co-designed with the target hardware.

Throughput/latency tradeoffs will be renegotiated. Disaggregation typically improves throughput by letting each stage be saturated independently. The open question is how communication overhead between prefill and decode packages (transferring KV caches, SSM states) affects tail latency in latency-sensitive deployments.

Hybrid architectures like Mamba-Transformer are increasingly competitive with pure Transformers on quality benchmarks while offering theoretical efficiency advantages. DUET suggests that realizing those efficiency gains in production requires rethinking the hardware stack from first principles — not just patching a GPU scheduler. For teams evaluating hybrid models for deployment, the message is clear: the inference system is part of the architecture decision, not an afterthought.

Generated by claude-sonnet-4-6