GPU Schedulers Are Blind to AI Agents — Here's the Fix
I don’t have WebFetch access, so I’ll write the explainer from the abstract and my knowledge of this problem space.
The Mismatch at the Heart of Agentic AI Infrastructure
When you deploy an LLM-powered agent today — one that plans, retrieves, reflects, and acts over multiple steps — your GPU cluster has no idea it’s doing any of that. As far as the scheduler is concerned, each individual inference call is an anonymous, stateless HTTP request. The fact that call #47 is semantically continuous with calls #1 through #46, sharing context, KV cache, and intermediate reasoning state, is invisible to the infrastructure beneath it.
This abstraction mismatch is the core problem that SAGA attacks. And the cost is not subtle: the authors report 3–8x end-to-end latency inflation for compound AI workloads compared to what the compute would theoretically allow.
Why Request-Level Scheduling Breaks for Agents
Modern LLM serving systems like vLLM, TensorRT-LLM, and SGLang are optimized around the request as the unit of work. A request arrives, gets batched with others, runs inference, returns a response, and the system moves on. KV cache is evicted or managed purely on a per-request basis with no knowledge of whether the caller intends to make nine more related calls immediately after.
Agent workflows violate every assumption baked into this model:
They’re long-running and chained. A ReAct-style agent loop or a multi-step tool-using pipeline might execute 50–200 sequential LLM calls. Each step’s output is the next step’s input. The inter-call latency — queuing, preemption, cold-start cache misses — accumulates multiplicatively.
They carry expensive intermediate state. The KV cache generated during early reasoning steps is directly reusable by later steps in the same workflow. But if that cache gets evicted between calls (because the scheduler sees idle capacity and reclaims it), the next call recomputes it from scratch. At model scale, this is gigabytes of wasted computation per agent task.
They have workflow-level SLOs, not request-level ones. A user waiting for an agent to complete a task doesn’t care how fast any individual inference step is — they care about total wall-clock time. Optimizing request-level throughput while ignoring cross-request coherence is optimizing the wrong metric.
SAGA’s Approach: Program-Level Scheduling
SAGA reframes the schedulable unit from the individual inference call to the entire agent workflow — what the authors call program-level scheduling. The workflow becomes a first-class citizen in the distributed system, with the scheduler maintaining awareness of its full execution graph rather than seeing only individual requests as they arrive.
The key insight is that agent workflows have predictable structure. A planning agent will generate a plan, then execute tool calls, then synthesize results. A RAG pipeline will embed, retrieve, then generate. These patterns can be modeled as DAGs, and a workflow-aware scheduler can make radically better decisions than a stateless request queue.
Concretely, SAGA appears to provide:
Workflow-atomic placement. Rather than scattering an agent’s sequential calls across different GPU workers based on momentary load, the system can colocate them — keeping KV cache warm and avoiding the serialization overhead of shipping state across nodes between steps.
Cross-step cache management. Because the scheduler knows call N+1 is coming from the same workflow as call N, it can make informed decisions about KV cache retention rather than treating eviction as purely a local memory-pressure problem.
Distributed coordination. The “distributed” qualifier in the abstract is significant — SAGA is designed for multi-GPU, multi-node clusters, not just single-server optimization. Workflow state needs to be tracked and routed correctly across the cluster as requests flow through the system.
Putting Numbers to the Problem
The 3–8x latency inflation figure in the abstract is the headline claim, and it’s plausible given how the inefficiencies stack. Consider an agent making 100 sequential calls where each call incurs 50ms of additional latency from queuing, preemption, and cache misses compared to an idealized warm-cache scenario. That’s 5 extra seconds of pure scheduling overhead, on top of whatever the actual compute takes. At longer workflows or larger models, the multiplier gets worse.
The gigabytes of intermediate state figure is similarly concrete. A 70B parameter model with a 32k token context generates KV cache on the order of several GB. If an agent’s early context — its system prompt, initial reasoning, retrieved documents — gets evicted and recomputed at every step, the wasted memory bandwidth alone is significant before you even count the compute cycles.
What This Means in Practice
SAGA represents a bet that the right place to fix agentic latency is the infrastructure layer, not the application layer. Developers currently work around these problems with techniques like prompt caching (Anthropic, OpenAI), explicit KV cache pinning (some vLLM configs), or batching agent steps where possible — all application-level hacks that patch around a scheduler that can’t see the full picture.
A workflow-aware scheduler would let those optimizations happen automatically and more aggressively, without requiring developers to hand-tune their agent loop. It also opens the door to smarter preemption policies: if the scheduler knows a workflow has been running for 80 steps and is likely near completion, it can deprioritize preemption differently than it would for a cold request.
Watch for SAGA’s benchmarks against vLLM and similar systems on multi-step agent workloads — the degree to which the gains hold across different workflow topologies (linear chains vs. branching tool use vs. parallel subagents) will determine how general the approach is. The harder question is adoption: retrofitting workflow-level semantics into existing serving infrastructure requires API and protocol changes that span the full stack from orchestration frameworks down to the GPU scheduler.