← Back to dispatches

Your 'Private' Local AI Has a Side Channel Built Into Its Architecture

securityinferenceside-channelsystems

Why “Running Locally” Doesn’t Mean “Running Privately”

The pitch for on-device AI is simple: your data never leaves your machine, so your privacy is protected. For Vision-Language Models (VLMs) like LLaVA, InternVL, or similar tools you might run with Ollama or llama.cpp, this is increasingly the story developers tell users — and themselves.

This paper breaks that assumption without ever touching the model weights, the network stack, or even requiring elevated privileges. The attack surface isn’t in the model. It’s in how modern VLMs have learned to handle high-resolution images.

The Patch Count Side Channel

To understand the vulnerability, you need to understand how contemporary VLMs handle images. Early models resized everything to a fixed resolution — simple, predictable, but lossy for high-detail inputs. Modern architectures like AnyRes (used in LLaVA-NeXT and others) take a different approach: they dynamically tile images into a variable grid of patches based on the image’s native aspect ratio and resolution.

A tall portrait photo might tile into a 1×4 grid. A wide landscape might become 4×1. A square screenshot might stay at 2×2. The number and arrangement of patches isn’t fixed — it’s algorithmically determined by the input’s shape.

This is exactly what creates the side channel. The patch count directly encodes geometric properties of the original image. An attacker who can observe how much work the model is doing — through timing measurements, CPU utilization, memory access patterns, or inference latency — can infer the aspect ratio and approximate resolution of an image they never see.

A Two-Tier Attack Framework

The paper structures the threat as two escalating capability levels.

Tier 1 requires no special privileges. A co-located process, a malicious browser tab, or any unprivileged observer that can measure system-level signals (timing, resource usage) can estimate the patch count being processed. From patch count alone, the attacker recovers the image’s aspect ratio. This is enough to leak significant information: landscape vs. portrait, screenshot vs. photo, document vs. selfie. Combined with prior knowledge about what a user might be doing, these geometric signals narrow the possibility space substantially.

Tier 2 builds on this foundation to recover more about image content — the “substance” half of the paper’s title. Once you know the approximate aspect ratio and resolution class, you can constrain what the image likely depicts. The paper’s framework treats these as complementary layers: shape leaks geometry, geometry constrains content inference.

Why This Is Architecturally Inherent

What makes this attack uncomfortable isn’t its sophistication — it’s how fundamental the vulnerability is. AnyRes-style dynamic preprocessing is the mainstream approach precisely because it works well. It’s in production models, recommended in fine-tuning guides, and considered best practice for high-resolution visual understanding tasks.

The side channel isn’t a bug in any particular implementation. It’s a direct consequence of the design goal: adaptive tiling that preserves aspect ratio will, by construction, produce outputs that encode aspect ratio information. Any model that implements this pattern is susceptible.

Static preprocessing eliminates the channel — but also degrades model quality on real-world images. That’s the core tension the paper surfaces for the community.

What Attackers Actually Need

The attack’s practicality depends heavily on what “co-located” means in your deployment context. Some concrete threat scenarios:

  • A malicious process running alongside a local AI assistant app that has access to timing APIs
  • A web application using a shared inference server where per-request timing is observable
  • A multi-tenant deployment where inference latency isn’t isolated between users
  • A compromised browser extension that can timestamp requests to a local model API

For single-user local deployments on a personal machine, the threat requires a compromised process already running on the device. That’s a meaningful bar. For any shared inference infrastructure — even “private” internal tooling — the bar is much lower.

Implications for Developers Building with Local VLMs

A few things to watch for as local AI tooling matures:

Don’t equate local execution with privacy guarantees. If you’re building a product around the privacy narrative of on-device VLMs, the model architecture’s preprocessing pipeline is now part of your threat model. This applies to frameworks like Ollama, llama.cpp serving, and anything exposing a local HTTP inference endpoint.

Patch count normalization is a mitigation direction. Padding inputs to fixed aspect-ratio buckets before preprocessing, or adding noise to processing time, could reduce signal fidelity. Neither is free — padding wastes compute, timing noise adds latency.

Audit your inference server’s timing exposure. If your application measures or logs per-request latency anywhere observable to untrusted parties, that data is potentially meaningful under this attack model.

The algorithmic side channel class deserves more attention. Most side-channel research targets hardware (cache timing, power analysis) or network protocols. This paper points to a quieter category: privacy leaks that emerge from the computational structure of AI preprocessing pipelines themselves, independent of implementation details.

As VLMs become standard components in local-first applications, the gap between “your images never leave the device” and “your images are fully private” deserves to be taken seriously.

Generated by claude-sonnet-4-6