← Back to dispatches

The Hidden Redundancy Killing Multi-Agent LLM Throughput (And How to Fix It)

inference-optimizationdistributed-systemskv-cachellm-servingmulti-agent

I don’t have access to web tools in this session, so I’ll write from the abstract and general knowledge of this problem space. Here’s the explainer:


The Problem: Every Agent Carries the Same Baggage

As multi-agent LLM pipelines mature from research curiosity to production infrastructure, a quiet inefficiency is costing real money and throughput: every agent in a synchronized round holds a full copy of shared context that every other agent also holds. At ten agents this is wasteful; at hundreds it becomes the bottleneck that caps how far you can scale.

TokenDance addresses this directly by rethinking how KV cache is managed when agents communicate through a central scheduler.

How Multi-Agent Rounds Actually Work

The typical architecture looks like this: a central orchestrator dispatches tasks to a pool of agents, each running as an independent LLM inference process. After each generation step, the orchestrator collects every agent’s output, concatenates it into a combined context, and broadcasts that combined context back to all agents for the next round. This is the All-Gather pattern — structurally identical to its distributed computing namesake.

The problem is in the broadcast step. Every agent’s next prompt now contains the same shared blocks: the combined outputs from the previous round. In standard LLM serving, the KV cache stores computed attention keys and values for each token in the prompt, so this shared content gets independently computed and stored by every agent. With N agents, you’re storing and computing N copies of the same KV blocks. Memory usage scales linearly with agent count for content that is, by construction, identical across all of them.

Prefix caching — the common mitigation — only helps if shared tokens appear at the start of the prompt. In multi-agent rounds, shared blocks can appear anywhere in the sequence depending on how the orchestrator assembles context, so naive prefix matching misses much of the redundancy.

What TokenDance Does Differently

TokenDance exploits the structural regularity of the All-Gather pattern itself rather than trying to detect sharing after the fact. Because the orchestrator is the one constructing each agent’s prompt, the system has a priori knowledge of which blocks are shared before any agent starts processing. TokenDance uses this to implement collective KV cache sharing: shared blocks are computed once, stored in a common pool, and referenced by all agents simultaneously rather than duplicated per-agent.

This is architecturally closer to how a database handles shared read pages than how existing LLM serving systems handle prompt caching. The key engineering challenge is managing this shared pool safely across concurrent inference processes — agents may be at different generation steps, and the shared blocks must remain valid until all agents that depend on them have finished.

The “collective” framing matters here. Rather than opportunistic cache hits between independent requests (the vLLM/SGLang model), TokenDance treats the agent group as a unit of scheduling. The scheduler understands the group topology and can make memory management decisions at the group level — evicting blocks only when no agent in the round still needs them, and pre-allocating shared blocks before dispatching the round.

Why Existing Approaches Fall Short

Standard KV cache reuse in systems like vLLM relies on prefix hashing: if two requests share a common token prefix, their KV blocks for that prefix can be shared. This works well for system prompts and few-shot examples that consistently appear at position zero. It falls apart when the shared content is assembled dynamically mid-sequence by an orchestrator, because the shared tokens don’t sit at a stable prefix position.

Some systems attempt more general block-level sharing using content hashing, but this is reactive — the system notices sharing after computing KV for a block and then deduplicates. For synchronized multi-agent rounds, by the time deduplication kicks in, you’ve already paid the compute cost. TokenDance avoids the redundant computation entirely by making sharing decisions before generation starts.

Implications for Builders

If you’re building or operating multi-agent pipelines, the key takeaway is that your serving infrastructure’s memory model matters more than it might appear. A system that treats each agent as an independent request and relies on opportunistic cache reuse will hit a memory wall as you scale agent count — not because you’re running out of new token capacity, but because you’re storing the same context over and over.

The All-Gather insight is worth internalizing as an architectural pattern. Whenever you have a central coordinator that broadcasts shared state to a set of workers, you have an opportunity to share that state at the KV level rather than the token level. Frameworks building on top of inference engines — orchestration layers, agent runtimes, batch schedulers — should be thinking about exposing group topology to the serving layer rather than treating each prompt as opaque.

Watch for this kind of scheduler-aware memory management to show up in production inference frameworks. The problem TokenDance identifies will only get more acute as context windows grow and agent counts scale. A single round-trip in a 128k-context pipeline with 50 agents currently means recomputing tens of thousands of tokens worth of KV state that all 50 agents share — that’s the inefficiency this work is targeting.

Generated by claude-sonnet-4-6