Your LLM Serving Optimization Is Also a Spy: The Prefix Caching Side-Channel
When you’re running an LLM inference service for multiple customers, you’re doing more than sharing compute — you may be inadvertently sharing secrets. A new paper, CacheSolidarity, demonstrates that Automatic Prefix Caching, one of the most effective optimizations in modern LLM serving, opens a measurable timing side channel that allows tenants to probe each other’s request history.
The Optimization That Enables the Attack
To understand the vulnerability, you need to understand what Automatic Prefix Caching (APC) actually does. LLM inference is dominated by the attention mechanism, which requires computing key-value (KV) pairs for every token in the context. For long or repeated prefixes — system prompts, RAG retrieval templates, few-shot examples — recomputing these KV pairs on every request is wasteful. APC solves this by hashing the token sequence of a prefix and storing the resulting KV cache. If a subsequent request starts with the same tokens, the system reuses the cached state, skipping the prefill computation entirely.
The latency difference is real and substantial. A cache hit on a long system prompt can shave hundreds of milliseconds off time-to-first-token (TTFT). In a high-throughput serving system like vLLM or SGLang, this is a meaningful throughput gain — which is exactly why every major inference framework has adopted it.
Turning a Performance Feature Into a Spy
The attack logic is straightforward once you see it. In a multi-tenant deployment, tenants share the same KV cache pool. An attacker controls one tenant. They can send requests with arbitrary prefixes and precisely measure TTFT for each. A fast response indicates a cache hit — meaning someone else’s request already populated the cache with that prefix.
This translates directly into an information leak. Consider a few concrete scenarios:
System prompt extraction. Many API providers differentiate their products through carefully crafted system prompts that are kept confidential. An attacker can enumerate candidate prefixes — trying variations of common prompt patterns — and use cache-hit timing as an oracle. Sequences that produce a fast TTFT were recently seen by the system, meaning they’re likely part of another tenant’s system prompt.
User activity inference. If a platform caches user session context, an attacker sharing the inference backend can probe whether a specific user is actively making requests by testing whether their known session prefix produces a cache hit.
Membership inference. In RAG systems, retrieved document chunks are often prepended as context. An attacker with access to the same document corpus can determine which documents were recently retrieved by other users — effectively leaking query patterns and user behavior.
The attack requires no special access. It’s entirely achievable by a legitimate API customer making ordinary requests and measuring response latency.
What CacheSolidarity Proposes
The paper’s defense, CacheSolidarity, is built around the observation that the side channel exists because tenants share a single undifferentiated cache namespace. The core insight is to enforce cache isolation without abandoning the throughput benefits of prefix reuse.
The approach segments the KV cache by tenant identity, ensuring that a cache entry populated by tenant A cannot produce a hit for tenant B. This eliminates cross-tenant leakage by construction. The challenge is doing this without destroying the efficiency gains — if every tenant has a fully independent cache, you lose the deduplication benefit for genuinely shared content (common system prompts, shared instruction formats).
CacheSolidarity addresses this with a tiered model. A public cache layer holds entries for content that is provably non-sensitive: prompts that have been explicitly marked as shareable, or common instruction prefixes with no confidential content. A private cache layer is strictly per-tenant. The framework provides mechanisms for tenants to declare prefix segments as shareable, allowing the serving system to route them to the public tier while keeping proprietary prompt content isolated.
The paper also addresses timing normalization as a complementary defense — padding response latency to reduce the signal-to-noise ratio of timing measurements — though this comes with throughput tradeoffs.
Why This Matters for Inference Infrastructure
This class of attack isn’t new in principle — cache timing side channels have existed in CPU and memory systems since at least Flush+Reload. What’s notable here is that the LLM serving stack has recreated the same vulnerability pattern independently, and the ecosystem has been deploying APC broadly without explicit security analysis of the multi-tenant case.
If you’re running a shared inference service — whether a commercial API, an internal platform for multiple teams, or a model-as-a-service product — a few things warrant attention:
Audit your caching scope. Most frameworks enable APC globally. If tenants are sharing a cache backend, you have this exposure today.
Treat system prompts as secrets at the infrastructure level. Even if your application layer treats them as confidential, your serving layer may be exposing them through timing.
Watch for per-tenant cache partitioning features. vLLM and SGLang will likely need to add first-class support for tenant-scoped cache namespaces. CacheSolidarity provides a design reference for what that should look like.
Consider TTFT jitter for sensitive deployments. Adding controlled noise to response latency is a blunt instrument but an immediately deployable mitigation while framework-level fixes mature.
The broader lesson is that LLM serving infrastructure is now complex enough to harbor the same categories of vulnerabilities we’ve spent decades managing in operating systems and databases. Security analysis needs to keep pace with the optimization work.