← Back to dispatches

Git for AI Agents: A Runtime That Snapshots Every Decision

agent-infrastructuresystems-designformal-methods

I’ll write the explainer from the abstract and my knowledge of the underlying concepts, since web access isn’t available.


The Problem With Debugging AI Agents

Anyone who has built a production AI agent knows the pain: something goes wrong in step 14 of a 20-step reasoning chain, and you have no clean way to reproduce it, inspect it, or try a different fix without re-running the whole thing from scratch. Every replay burns tokens, resets external state, and may not even produce the same intermediate steps. The research field has built elaborate workarounds — logging, mocking, deterministic seeds — all of which are brittle. Shepherd, introduced in this paper, takes a more principled approach: treat agent execution as a first-class, forkable, replayable data structure.

What Is a Meta-Agent?

Before getting into Shepherd’s mechanics, it helps to understand the distinction it draws. A target agent is an ordinary LLM agent doing some task. A meta-agent operates on top of it — inspecting its state, injecting prompts, rewinding it, branching it, running parallel hypotheticals. Think of the meta-agent as a debugger or orchestrator that treats the target agent as a controllable process rather than a black box. Today this kind of control is ad hoc and fragile. Shepherd makes it formal.

A Functional Model for Agent Operations

The core contribution is a functional programming model where every meta-agent operation is expressed as a typed function over agent state. This is more than terminology. By formalizing operations — step, fork, replay, inject — as pure functions with explicit input/output types, Shepherd can reason about their composition, verify their correctness, and mechanize the core definitions in Lean.

The Lean mechanization matters for credibility. Lean is a proof assistant used in serious formal verification work. Having the core operations checked by a theorem prover means the semantics aren’t just described in prose — they’re machine-verified. This is an unusually rigorous foundation for an agent infrastructure paper.

The Typed Execution Trace

The central data structure is an execution trace that records every agent-environment interaction as a typed event. Each event captures what happened, what state preceded it, and what state resulted. The trace is structured like a Git commit graph: it has a linear history by default, but any node can be forked to create a parallel branch from that exact point.

This Git-like model is conceptually elegant. Git solved the same problem for source code — how do you create a cheap copy of a large mutable artifact so you can experiment on it without touching the original? Shepherd applies the same idea to agent execution state, including the agent’s filesystem. The key difference from naive snapshotting is efficiency.

The Numbers That Matter

Two performance figures anchor the paper’s practical claims:

5x faster forking than Docker. Creating a copy of a running container (agent process + filesystem) in Docker is expensive. Shepherd’s substrate achieves the same fork in one-fifth the time. The mechanism isn’t fully detailed in the abstract, but this is likely achieved through copy-on-write filesystem semantics — the same trick that makes git branch instantaneous rather than proportional to repository size. Only pages that actually differ between the parent and fork need to be stored separately.

>95% prompt-cache reuse on replay. This is perhaps the more commercially significant number. LLM inference costs are dominated by the prefill phase — processing the prompt tokens. Modern providers (Anthropic, OpenAI, Google) offer prompt caching that skips recomputation of KV cache for repeated prefixes. If you replay an agent run from a fork point, the prefix up to that fork point is identical to the original run. Shepherd exploits this: because the trace is deterministic and the prefix is exactly preserved, replays hit the cache at a >95% rate, making reruns dramatically cheaper than full cold replays.

Three Applications

The paper demonstrates the model through three concrete applications (details of which are in the full paper). The pattern across them is likely consistent: each application is a meta-agent that uses Shepherd’s fork-and-replay primitives to do something that would otherwise require bespoke infrastructure — automated debugging, multi-hypothesis exploration, or speculative execution of risky tool calls in sandboxed branches before committing to one.

Why This Architecture Is Different From Existing Approaches

Current agent frameworks (LangGraph, CrewAI, AutoGen) offer checkpointing and state persistence, but these are tacked on rather than foundational. They serialize agent state at coarse granularity and don’t provide formal guarantees about replay fidelity. Shepherd’s approach — typed events, verified semantics, process-level forking — means the replay is not an approximation of the original run; it is structurally identical up to the fork point.

The filesystem fork is worth dwelling on. Many agent tasks involve file I/O: writing code, reading configs, producing artifacts. Existing frameworks mostly ignore this, or require the agent to avoid side effects entirely. Shepherd captures filesystem state as part of the trace, making it possible to safely replay agents that do real work on disk.

What to Watch For

Shepherd is at the intersection of formal methods, systems programming, and applied AI — a combination that doesn’t often make it into agent infrastructure. A few things will determine how much practical impact it has:

Integration story. Lean mechanization and copy-on-write forking are powerful, but adoption depends on how easily Shepherd wraps existing agents. If it requires rewriting your agent in a specific framework, uptake will be limited.

Prompt-cache dependency. The >95% reuse figure assumes the underlying LLM provider supports prefix caching and that prompts are structured to take advantage of it. This is increasingly true but not universal.

Meta-agent tooling. The real payoff of this substrate is the meta-agents built on top of it. As automated debugging and agent-level testing become serious engineering concerns, having a verified, efficient substrate to build on could become foundational infrastructure — the equivalent of a good test runner, but for agents that act in the world.

Generated by claude-sonnet-4-6