← Back to dispatches

Stack-Spanning Attacks: How Compound AI Systems Create New Hardware-Software Attack Surfaces

securityAI-systemsRAGdistributed-systems

As I don’t have web access enabled, I’ll write the explainer based on the title, abstract, and my knowledge of this research domain.


When AI Pipelines Inherit Every Bug You’ve Ever Patched

Modern AI applications don’t run as isolated language models. They run as compound systems — chains of LLMs, vector databases, code interpreters, web browsing tools, and API integrations stitched together into autonomous pipelines. This architectural complexity is what makes them powerful, and it’s exactly what makes them a new class of attack surface.

The Cascade paper formalizes something the security community has been circling: when you build an AI system on top of a traditional software stack running on distributed hardware, you don’t just inherit the capabilities of that stack — you inherit all of its historical vulnerabilities too. And unlike patching a standalone web service, the attack paths through compound AI systems can chain software flaws with hardware-level exploits in ways that dramatically amplify the threat.

What “Gadgets” Means Here

The paper borrows the concept of attack gadgets from the world of Return-Oriented Programming (ROP), a well-established exploitation technique where attackers chain together small snippets of existing code — each individually harmless — to build a full exploit payload. Cascade applies this framing to compound AI systems: each component in the pipeline (an LLM, a retrieval layer, a code execution sandbox, a tool-calling interface) represents a potential gadget. None of them need to be individually catastrophic. The threat emerges from composition.

A compound AI pipeline might look like: user prompt → LLM reasoning layer → RAG retrieval against a vector DB → LLM synthesis → code interpreter → external API call. At each boundary, there are trust assumptions. Cascade examines what happens when those assumptions break, and — critically — how an attacker can exploit weaknesses at one layer to reach vulnerabilities in another.

The Software Attack Surface

The software side of the threat model draws directly from CVE-documented vulnerabilities in the components compound AI systems depend on. Vector databases and embedding stores often expose REST APIs with weak authentication or injection-susceptible query interfaces. Code interpreter sandboxes — increasingly common in agentic frameworks — run real code in environments that have historically been difficult to fully isolate. Tool-calling mechanisms frequently pass LLM-generated strings into shell commands, HTTP clients, or SQL interfaces with insufficient sanitization.

Prompt injection remains the entry point of choice. An adversarially crafted document retrieved via RAG can instruct the LLM to misuse a connected tool — trigger an SSRF via a fetching tool, exfiltrate data through a logging interface, or escalate into a code execution step with attacker-controlled input. The compound architecture means the LLM’s instruction-following capability becomes an amplifier for every downstream vulnerability in its toolchain.

The Hardware Layer

What distinguishes Cascade from prior compound AI security work is its explicit inclusion of the hardware infrastructure as an attack surface. Compound AI systems run on distributed compute — typically shared GPU or accelerator clusters, cloud inference endpoints, and commodity server hardware. This infrastructure is subject to the same hardware-level attacks that have plagued cloud computing: side-channel attacks on shared processors, Rowhammer-style DRAM exploits, and speculative execution vulnerabilities.

The key insight is that hardware-level access, once treated as a separate threat model from application-layer exploits, becomes reachable through the software attack gadgets described above. A code execution vulnerability in a sandboxed tool can expose the underlying host. A compromised inference node in a distributed pipeline can leak weights, intermediate activations, or — in multi-tenant deployments — data from adjacent workloads. The attack isn’t just horizontal across the AI pipeline; it’s vertical through the software-hardware stack.

Threat Amplification in Practice

The “amplification” framing is the paper’s central contribution. Traditional security analysis treats vulnerabilities in isolation: this CVE has a CVSS score, this prompt injection technique has a known impact. Cascade argues that in compound AI systems, the effective severity of a low-impact vulnerability in one component is a function of what it can reach in the rest of the pipeline.

A medium-severity server-side request forgery in a document retrieval service might be unremarkable on its own. Combined with a prompt injection vector and a code execution step downstream, the same vulnerability can serve as an initial access primitive into the compute infrastructure. The paper’s framework — “composing gadgets” — is essentially a methodology for mapping these amplification paths systematically, analogous to how exploit developers map ROP chains in binary exploitation.

What to Watch For

For teams building or securing compound AI systems, the practical implications fall into a few categories:

Trust boundaries matter more than component-level hardening. Patching the vector DB doesn’t help if the LLM will still forward attacker instructions to it. Treat the LLM’s instruction-following as an attacker-controlled input to every downstream tool.

Sandboxing is load-bearing infrastructure. Code interpreters and tool execution environments aren’t optional safety measures — they’re security boundaries. Their isolation properties need to hold against hardware-level attacks, not just application-layer escapes.

Distributed compute surfaces are in scope. Security reviews of AI pipelines shouldn’t stop at the application layer. The inference infrastructure, shared accelerators, and multi-tenant deployment patterns all represent gadgets that a sufficiently positioned attacker can incorporate into a chain.

The broader implication is that compound AI systems require a systems security methodology — one that models how vulnerabilities compose across layers — rather than checklist-based component auditing. As these architectures become the default pattern for deploying capable AI, the attack surface they present will only grow more complex.

Generated by claude-sonnet-4-6