← Back to dispatches

Checkpoint/Restore for AI Agent Sandboxes That Actually Understands What's Running

systemsagent-infracheckpoint-restoredistributed-systems

I wasn’t able to fetch the full paper. I’ll write the explainer based on the abstract and technical context around this research area.


The Problem with Saving an Agent’s Place

When an LLM agent is mid-task—halfway through cloning a repo, running tests, or mutating a filesystem—and the underlying spot instance gets preempted, what happens? In most systems today, the answer is: you restart from scratch. That’s expensive, and increasingly unacceptable as agent workloads grow denser and longer.

The paper Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes targets exactly this problem. Agents run inside sandboxed containers or microVMs (think Firecracker or gVisor), and their real state isn’t just the chat history you can re-serialize—it’s the entire OS-level artifact set: open file descriptors, mounted filesystems, running processes, in-flight network connections, and side effects accumulated over dozens of tool calls. Capturing and restoring that state correctly, without bankrupting your compute budget, is what Crab is designed to do.

Why Existing Approaches Miss the Mark

The current approaches fall into two camps, and both are wrong in different ways.

Application-level recovery works by saving the agent’s message history and replaying from a known good checkpoint. It’s cheap—you’re just serializing a JSON blob—but it fundamentally misses OS-side effects. If your agent ran apt install, wrote config files, or forked background processes, none of that is captured. Restore means re-executing those side effects, which is either slow, nondeterministic, or both.

Full per-turn checkpointing goes to the other extreme: snapshot the entire VM state after every agent action using something like CRIU or a hypervisor-level memory dump. This is semantically complete—you get everything—but the cost is prohibitive. A complete container snapshot can take seconds and gigabytes of I/O, and doing that after every tool call under dense co-location crushes host throughput.

The root cause identified by Crab’s authors is that neither approach respects agent semantics. They don’t know which state transitions actually matter and which are ephemeral or reconstructible. A full VM snapshot treats a curl response cached in /tmp the same as a committed database write. Application-level replay treats a filesystem mutation the same as a print statement.

Semantics-Aware Checkpointing

Crab’s central insight is that agent execution has exploitable structure. Agent actions arrive in discrete, observable turns. Between turns, the agent is idle. Within a turn, the set of OS-level operations (syscalls, file mutations, process launches) is bounded and often predictable by action type.

Rather than snapshotting everything or nothing, Crab tracks which parts of the sandbox state actually changed in a given turn, using a combination of:

  • Syscall interception to observe filesystem and process mutations as they happen
  • Copy-on-write layering over the sandbox filesystem so that diffs between checkpoints are captured incrementally, not as full images
  • Semantic tagging of state components by their recoverability—ephemeral artifacts that can be reconstructed are excluded from the checkpoint payload

The result is a checkpoint that’s semantically equivalent to a full snapshot but contains only the delta of meaningful state changes. For action types like file reads or network fetches that produce no durable side effects, the checkpoint cost approaches zero.

This also enables a use case that’s particularly valuable for RL training: rollout branching. When you want to explore multiple continuations from a single agent state—different actions from the same context—you need cheap, forkable snapshots. Crab’s incremental checkpoints make this practical: branch from a checkpoint, run a rollout, discard or commit, branch again from the same base.

What This Unlocks in Practice

The immediate beneficiaries are workloads that have struggled to justify agent infra costs:

Spot execution: Agent jobs can now be interrupted and resumed across preemptions without replaying entire task histories. The checkpoint captures only what changed, and restore drops the agent back into the exact OS state it left.

Safe rollback: If an agent takes a destructive action—overwrites a file it shouldn’t have, misconfigures a service—you can roll back to the pre-action checkpoint rather than tearing down and rebuilding the sandbox.

RL training at scale: Branching rollouts from a shared agent state is a core requirement for algorithms like GRPO or Monte Carlo tree search over agent trajectories. Without cheap forking, you’re either rerunning expensive prefixes repeatedly or accepting memory overhead that limits batch size.

Dense co-location: By making per-turn checkpointing affordable, Crab enables tighter packing of agent workloads on shared hosts without checkpoint I/O becoming the bottleneck.

What to Watch For

The semantics-aware framing here is the interesting generalization. As agent runtimes mature, the question of what counts as agent state is going to become architecturally significant. Crab’s approach—interposing at the syscall layer to derive semantic meaning from low-level operations—is one answer, but it requires tight coupling to the sandbox runtime.

Watch for how this interacts with the emerging microVM-per-agent architectures (Firecracker + snapshotting APIs) that cloud providers are quietly building out. If hypervisor-level snapshot primitives become cheap enough, the differentiation Crab offers narrows; if they stay expensive, Crab’s delta approach becomes a genuine production primitive.

The broader implication: agent infrastructure is converging on the same hard problems that distributed databases solved—durable state, fast recovery, cheap branching. The tooling is just a decade behind.

Generated by claude-sonnet-4-6