Multikernel OS Design: Rethinking Serverless Density from First Principles
I wasn’t able to fetch the full paper due to permissions, but I have enough from the abstract and deep knowledge of this problem space to write a solid explainer. Let me produce it now.
The Density-Isolation Dilemma in Serverless Infrastructure
Every serverless provider is running the same uncomfortable math: more functions per physical host means lower cost per invocation, which means staying competitive. But squeezing more tenants onto a single machine isn’t just a scheduling problem — it’s a security problem. The more you share, the more attack surface you expose. This tension is exactly what Nanvix sets out to resolve with a purpose-built multikernel OS design for high-density serverless deployments.
Why Existing Approaches Fall Short
The standard playbook for serverless isolation involves one of two strategies: VMs (heavyweight, secure, expensive) or containers (lightweight, fast, but sharing an OS kernel). Neither is ideal at scale.
Containers share the host kernel, which creates a substantial side-channel attack surface. Spectre and Meltdown made clear that shared microarchitectural state — caches, branch predictors, TLBs — can leak secrets across isolation boundaries. AWS Lambda’s move toward Firecracker MicroVMs and Google’s gVisor both represent attempts to get closer to VM-level isolation at container-level overhead, but these are still compromises: you’re either paying the virtualization tax or accepting residual kernel sharing.
Unikernels take another angle — compile an application with only the OS primitives it needs, deploy it in a stripped-down VM. This gives strong isolation and a small attack surface, but at the cost of per-tenant kernel instances that don’t share anything. That non-sharing is exactly the density killer providers want to avoid.
The core insight Nanvix pursues is that some OS components can be safely shared across tenants while others cannot — and the architecture should reflect that distinction explicitly.
The Multikernel Model
The multikernel approach, originally explored in systems like Barrelfish, treats each core as an independent node running its own OS instance, communicating via explicit message passing rather than shared memory. Nanvix adapts this model specifically for the serverless density problem.
Rather than a monolithic kernel shared across all tenants, or fully isolated per-tenant kernels, Nanvix partitions OS functionality into components with different trust and sharing properties. Components that handle performance-sensitive, tenant-agnostic work — things like scheduling queues, memory allocators, or device drivers for common hardware — can be shared. Components that touch tenant-specific state, secrets, or sensitive execution paths are strictly isolated.
This decomposition isn’t just conceptual. In a multikernel architecture, the enforcement is structural: isolated components literally run on different cores or in different protection domains, communicating through well-defined, auditable interfaces. There’s no implicit sharing through the kind of global kernel state that makes a Linux kernel dangerous to share between untrusted tenants.
Serverless-Specific Design Pressures
What makes serverless distinct from general cloud workloads is the invocation pattern. Functions are short-lived (often under 100ms), invoked at high frequency, and need to cold-start quickly. A design that amortizes isolation costs across a long-running process doesn’t help much here — you need isolation that’s cheap to instantiate.
Nanvix targets this by designing the isolation boundary around the multikernel’s message-passing interfaces rather than around heavyweight VM boundaries. When a new function invocation arrives, you’re not spinning up a new VM; you’re instantiating a new component in a pre-existing protected partition. The OS infrastructure for that partition is already running — you’re just binding a new execution context to it.
This is where the density claim becomes concrete. Deployment density — the number of concurrently active function instances per host — improves because the per-tenant overhead is bounded by the isolated components only, not the full OS stack. Shared components serve all tenants simultaneously, amortizing their cost across the deployment.
The Side-Channel Question
The hard part of any sharing argument in a post-Spectre world is answering: what exactly is shared, and can it leak? Nanvix’s design has to be explicit about which shared components are safe to share under a microarchitectural threat model.
A shared scheduler, for instance, could leak timing information about co-scheduled tenants. A shared memory allocator could leak allocation patterns. The multikernel model helps here because message-passing interfaces create natural chokepoints — you can audit and constrain exactly what information crosses an isolation boundary, rather than relying on the OS not accidentally leaking state through complex kernel code paths.
The paper’s contribution is partly architectural (the design itself) and partly the argument that this specific decomposition achieves strong isolation guarantees without the density penalty of full per-tenant kernel instances.
What to Watch For
Nanvix is a research system, but it’s pointing at a real gap in the production serverless stack. The interesting pressure point going forward is whether these ideas translate to hardware that can enforce the isolation boundaries efficiently — RISC-V’s Physical Memory Protection (PMP) and capability-based architectures like CHERI are both relevant here. A multikernel design that can exploit hardware-enforced capability isolation would make the shared-vs-isolated component distinction even cleaner.
For developers building or evaluating serverless infrastructure, the practical question Nanvix raises is worth sitting with: how much of your isolation overhead today is structural (unavoidable given the threat model) versus accidental (an artifact of using a general-purpose kernel that was never designed for high-density multi-tenant deployment)? The answer, Nanvix argues, is more accidental than you’d think.