One Fast Path for All: The Cloud Network Stack That Lets Tenants Define Their Own Protocols
I don’t have WebFetch access in this session, so I’ll write the explainer from the abstract and domain knowledge, clearly grounding claims in what the paper states.
The Hidden Tax of Cloud Networking
Every time a packet travels between two VMs in a cloud datacenter, it pays a layering toll most developers never think about. It leaves the guest application, descends through the guest kernel’s network stack, crosses into the hypervisor via a paravirtualized interface like virtio, climbs back up through the host kernel’s vhost subsystem, and finally reaches the physical NIC driver. That round-trip through guest and host layers isn’t free: it burns CPU cycles and, critically, inflates tail latency — the kind that makes P99 latency charts look like mountain ranges. For latency-sensitive workloads like distributed databases, RPC frameworks, or storage backends, this overhead is a genuine performance ceiling.
The cloud industry’s answer to this has been the shared host datapath: collapse all those layers into a single, optimized userspace process running on the host, handling packets for all tenants at once. Google’s Snap system, deployed at scale in its dataproduction, demonstrated that this model can dramatically reduce CPU cost and latency. But it came with a sharp tradeoff. These stacks are fixed-function: the cloud provider defines the protocols, and tenants get whatever they’re given.
Why Tenants Want Their Own Protocols
This matters more than it might seem. Modern distributed systems are deeply opinionated about their network behavior. A high-frequency trading workload needs radically different congestion control than a bulk-transfer backup job. A disaggregated storage system might want to implement custom reliability semantics directly in the transport layer to avoid the overhead of TCP’s general-purpose guarantees. RDMA-based applications want to bypass conventional transport entirely. The fixed-function constraint forces tenants to either accept a lowest-common-denominator protocol or pay the layering tax all over again by running their own stack inside the guest.
eBPF, which lets developers load verified programs into the Linux kernel at runtime, seems like the obvious escape hatch. The kernel already uses eBPF hooks to let operators customize packet processing in XDP and tc. Why not expose those hooks to tenants in a shared datapath?
The problem, as Chamelio identifies, is that eBPF has two structural weaknesses in this context. First, its programs are hook-sized: the extension points are narrow entry points, not a mechanism for implementing a full protocol stack. You can filter or mangle packets, but implementing a custom reliable transport with flow control, retransmission, and connection state is a different order of magnitude. Second, and more insidiously, the eBPF verifier provides safety, not performance isolation. It checks that a program won’t crash the kernel — no invalid memory accesses, no unbounded loops — but it says nothing about how long the program will take to run per packet. One tenant’s computationally expensive eBPF program can silently degrade throughput and latency for every other tenant sharing the datapath.
What Chamelio Does Differently
Chamelio is a shared cloud network stack designed to give tenants programmable protocol customization while preserving the performance isolation that makes shared datapaths valuable in the first place.
The system’s core insight is that programmability and isolation are not inherently in tension — they just require a different architectural contract than raw eBPF hooks provide. Rather than treating tenant code as event handlers bolted onto a fixed stack, Chamelio is built around the idea of tenant-defined protocol pipelines: tenants can supply code that implements substantial pieces of their network stack, but this code runs within a framework that enforces per-tenant resource budgets.
The name is telling: a chameleon changes its appearance while remaining the same animal. The shared infrastructure stays constant — the memory management, scheduling, NIC interaction, and multitenancy machinery — but the packet-processing logic can vary per tenant.
For performance isolation specifically, Chamelio needs to solve a problem the eBPF verifier deliberately ignores: bounding CPU time consumed by tenant code at runtime, not just proving absence of infinite loops statically. This requires a runtime accounting mechanism that the verifier-based model doesn’t provide.
Why This Is Hard to Build Well
The engineering challenge here is subtle. A shared datapath processes packets at very high rates — potentially millions per second across all tenants — so any per-packet overhead for scheduling, accounting, or context-switching between tenant code modules has to be vanishingly small. The whole point is to beat the latency and CPU cost of the layered virtualization model; if the isolation machinery eats back those gains, you’ve solved nothing.
This is the domain where systems papers earn their keep: the gap between “isolation is conceptually achievable” and “isolation is achievable without degrading the fast path” is filled with careful data structure choices, careful avoidance of synchronization on the critical path, and often some counterintuitive decisions about when to be lazy versus eager.
What to Watch For
Chamelio sits at an interesting intersection of several active areas: the push toward programmable network infrastructure (P4, SmartNICs, eBPF), the consolidation of cloud host software into shared userspace datapaths, and the ongoing tension between tenant flexibility and provider-enforced isolation guarantees.
For developers building latency-sensitive infrastructure — storage systems, RPC layers, distributed databases — the practical implication is that the dream of running custom transport protocols without paying a virtualization penalty may be getting closer to deployable reality. And for anyone building cloud infrastructure itself, the framing here — that performance isolation and programmability require a purpose-built runtime contract, not just a safety verifier — is likely to influence how the next generation of kernel and hypervisor extension points get designed.