← Back to dispatches

Generation IS Compression: Rethinking Video Codecs as Generative Latents

inferencegenerative-modelscompressionsystems

Video Compression Is Getting a Generative Overhaul

The video you stream tonight was probably encoded by a codec — H.264, H.265, AV1 — that works by discarding redundant information and transmitting what’s left. These codecs are engineering marvels, but they treat compression as a fundamentally destructive process: throw away bits you can reconstruct with math, keep everything else. The implicit assumption is that reconstruction means recovery, not generation.

A new paper, Generation Is Compression: Zero-Shot Video Coding via Stochastic Rectified Flow, challenges that assumption at its root. The core idea: if a powerful video generative model already “knows” what realistic video looks like, why not let it do the reconstruction? Better yet — why not make it the codec?

The Problem With Bolting Generators Onto Old Codecs

Previous generative compression approaches used neural networks as a kind of cosmetic layer. A conventional codec would encode and decode video the traditional way, then a generative model would be applied afterward to clean up compression artifacts or hallucinate missing detail. This is better than nothing, but it means you’re still paying the full cost of conventional coding, and the generative model is working against the grain — it’s trying to recover information the codec already threw away.

This paper takes a cleaner position: the generative model is the decoder. The encoder’s job is to transmit just enough information to steer the generation process toward the correct output.

Rectified Flow as a Coding Channel

The technical vehicle here is rectified flow, a generative modeling framework (used in models like Flux and Stable Video Diffusion 3) that maps noise to data along straight trajectories in latent space. Formally, a rectified flow model defines a deterministic ODE: given a noise initialization and a learned velocity field, you integrate from t=0 to t=1 and arrive at a generated sample.

The key insight is that this deterministic ODE can be rewritten as an equivalent stochastic differential equation (SDE). Why does that matter? Because SDEs have a well-known decomposition into a deterministic drift term (the signal) and a stochastic diffusion term (controllable noise). By encoding information into the noise schedule of this SDE rather than directly encoding pixel values, you can steer a pretrained generative model’s output with a compact bitstream.

In GVC, the encoder transmits a trajectory specification — essentially, quantized noise samples at carefully chosen timesteps along the generative path. The decoder, armed with a frozen pretrained video foundation model, integrates the SDE from that specification and recovers a video that matches the original. The transmitted bits don’t describe the video directly; they describe how to generate it.

Zero-Shot Is the Big Claim

The “zero-shot” claim deserves scrutiny, because it’s doing a lot of work. What the authors mean is that GVC requires no retraining, fine-tuning, or codec-specific training data. You take a pretrained video generation model off the shelf, apply the SDE reformulation, and it becomes a decoder. The encoder is a solver that inverts the generation process — finding the noise trajectory that would produce the target video.

This is non-trivial. Inversion in diffusion/flow models is an active research area, and doing it well enough to transmit only a compressed representation of the trajectory, while keeping reconstruction quality high, requires careful handling of quantization error propagation along the integration path.

The practical implication is significant: as video foundation models improve, GVC-style codecs get better for free. Drop in a newer, more capable generator, and compression quality improves with no additional engineering work.

What This Looks Like in Practice

At low bitrates — the regime where traditional codecs struggle most — generative approaches have a structural advantage. Conventional codecs break down at low bitrates because they’re forced to discard too much information, producing blurring and blocking artifacts. A generative decoder doesn’t recover missing information; it synthesizes plausible information. For smooth motions, natural textures, and predictable scenes, this can look substantially better, even if it isn’t pixel-accurate.

The tradeoff is fidelity for perceptual quality. GVC is not trying to achieve low MSE — it’s trying to produce video that looks correct to a human observer. This makes it well-suited for streaming, conferencing, and broadcast scenarios, less so for archival or forensic applications where exact reconstruction matters.

Implications for the Codec Ecosystem

A few things are worth watching as this line of work matures. First, the compute asymmetry: GVC decoding requires running a large video generation model, which is far more expensive than a traditional codec. Hardware acceleration for diffusion inference is improving rapidly, but real-time decoding at scale remains a challenge.

Second, the intellectual property questions around using foundation models as infrastructure are unresolved. If the codec is effectively a frozen copy of a proprietary generative model, deployment raises licensing questions that don’t exist for hand-engineered codecs.

Third, and most interesting from an engineering perspective: this architecture suggests a future where the “codec” and the “AI upscaler” are the same artifact. Rather than running a separate super-resolution or frame interpolation model on top of a conventional stream, generation-as-compression bakes that capability into the transmission protocol itself. That’s a meaningful architectural simplification, and it points toward a tighter coupling between content delivery and generative AI infrastructure than anything we have today.

Generated by claude-sonnet-4-6