← Back to dispatches

The Hidden Token Tax Killing Your Agentic Workflows

inference-optimizationagentic-systemsperformance-engineering

I don’t have WebFetch access, so I’ll write the explainer based on the abstract and my existing knowledge of MCP architecture and related LLM context research.


The Hidden Cost Buried in Every Tool Call

If you’ve built an agentic system on top of MCP — wiring up a handful of servers for file access, web search, database queries, and calendar tools — you’ve likely noticed that your prompts are enormous before a single user message even arrives. This paper names that phenomenon formally: the MCP Tax, and quantifies it at roughly 10,000 to 60,000 tokens of overhead per turn in typical multi-server deployments. That’s not a rounding error. At the high end, you may be spending the majority of your context window just describing tools the agent will never use in that turn.

The paper, Tool Attention Is All You Need, diagnoses the structural reason this happens and proposes two complementary mechanisms — Dynamic Tool Gating and Lazy Schema Loading — to eliminate it without sacrificing tool accessibility.

Why Eager Schema Injection Is the Root Problem

MCP’s current design is stateless by intent: each turn, every connected server dumps its full tool schema into the context. This is the “eager injection” model. An agent connected to five MCP servers, each exposing a dozen tools, might inject 60+ complete JSON schemas before the model sees the user’s message.

The costs compound in two ways. First, there’s the raw token overhead — every schema inflates the key-value (KV) cache that transformers maintain during inference, increasing memory pressure and latency. Second, and more insidiously, there’s what the paper calls reasoning degradation near fracture points. Prior work on long-context transformers has identified thresholds — typically somewhere between 60–80% context utilization depending on the model — where instruction-following and reasoning quality begin to measurably drop. A 60k-token tool payload on a 128k-context model puts you at roughly 47% utilization before the conversation starts. Add a few turns of history and tool outputs, and you’re in degradation territory.

This means the MCP Tax isn’t just expensive in dollars and latency — it actively makes your agent dumber.

Dynamic Tool Gating: Attend Before You Inject

The first proposed mechanism draws an analogy to the attention mechanism itself: rather than injecting all tool schemas unconditionally, gate which tools get injected based on a lightweight relevance signal computed from the incoming request.

The intuition is straightforward. A user asking “draft an email to the engineering team” doesn’t need the schema for execute_sql_query or read_file. Dynamic Tool Gating routes the request through a cheap scoring pass — potentially a smaller model, an embedding similarity check, or a learned classifier — to select a relevant subset of tools before context assembly. Only the schemas that clear the relevance threshold get injected.

The “attention” framing in the title is deliberate: this is applying selective attention at the orchestration layer rather than inside the transformer. The gating decision happens before the expensive model call, keeping the primary model’s context clean and focused.

Lazy Schema Loading: Just-in-Time Definitions

The second mechanism addresses a different failure mode: sometimes you don’t know which tools are needed until the model is mid-reasoning. Lazy Schema Loading defers full schema injection until the moment a tool is actually invoked.

In practice, this means the initial context contains only tool stubs — names and one-line descriptions, perhaps a few hundred tokens total — rather than complete JSON schemas with every parameter, type annotation, and example. When the model emits a tool call, the orchestration layer intercepts it, fetches the full schema for that specific tool, and re-presents it for argument generation.

This trades a small amount of latency (one extra round-trip per novel tool invocation) for a dramatic reduction in baseline context size. For workflows where an agent might browse 20 available tools but invoke only 2-3, the savings are substantial.

What the Numbers Suggest

The 10k–60k token range cited for the MCP Tax isn’t theoretical — it reflects real practitioner measurements from multi-server deployments. At current API pricing, 60k tokens of wasted overhead per turn across a high-volume agentic application represents meaningful cost. More importantly, if your agent runs long multi-turn tasks and you’re hitting context limits, tool schemas may be the silent culprit crowding out actual task context.

The paper positions these techniques as composable: Dynamic Tool Gating reduces the schema set proactively, while Lazy Schema Loading ensures that even the reduced set is only fully materialized on demand.

What to Watch For

The obvious engineering tradeoff is gating accuracy. A too-aggressive relevance filter that incorrectly excludes a needed tool causes the agent to either hallucinate a call or fail silently — potentially worse than the original bloat problem. The paper’s framing suggests the gating mechanism needs to be conservative enough to recall edge cases while still providing meaningful reduction.

There’s also a question of latency budget. Lazy schema loading adds a synchronous fetch in the hot path of tool invocation. For latency-sensitive applications, the KV cache savings need to outweigh the added round-trip.

Longer term, this work points toward a broader design principle: MCP servers shouldn’t be passive schema broadcasters. A richer protocol might let agents advertise intent, letting servers respond with contextually relevant subsets. As agentic frameworks mature and tool registries grow to hundreds of servers, the difference between eager and lazy schema injection will only widen. This paper arrives at a useful moment — just as MCP is becoming infrastructure rather than experiment.

Generated by claude-sonnet-4-6