Particle Filtering for Language Models: Ensembling LLMs with Sequential Monte Carlo
I can’t access the web in this session, so I’ll write the explainer from the abstract and my knowledge of this research area. Note that specific numbers or experiment details I can’t verify from the paper itself are omitted — I’ll work from what’s established in the abstract and the underlying methods.
Why Naive LLM Ensembling Breaks Down
If you’ve ever tried to squeeze better performance out of language models by averaging their outputs, you’ve probably noticed that “combine two models” is easier said than done. Classical ML ensembling — bagging, boosting, mixture of experts — works beautifully for fixed-dimension classifiers. Language generation is a different animal: outputs are sequences built token by token, and the interaction between choices compounds across hundreds of steps.
The failure mode of naïve aggregation is subtle but serious. Averaging next-token probability distributions from two models at each step sounds reasonable, but it’s equivalent to making a local decision at every position without any global accounting. If model A strongly prefers one path through the sequence space and model B prefers another, averaging their per-token logits produces something neither model actually endorses — an incoherent blend that can underperform both. The same problem plagues prompt ensembling: averaging token probabilities across different phrasings of a prompt doesn’t capture that each phrasing is expressing a different full-sequence belief.
This is the gap that Ensembling Language Models with Sequential Monte Carlo addresses.
SMC as a Decoding Framework
Sequential Monte Carlo is a classic inference technique from Bayesian statistics and robotics — think particle filters used in GPS localization or tracking problems. The core idea: instead of tracking a single hypothesis, maintain a population of candidate states (particles), propagate them forward, and periodically reweight and resample them based on how well they fit available evidence.
Applied to language model decoding, each particle is a partial sequence hypothesis. At every generation step, you extend each particle by sampling the next token, then reweight it based on scores from your ensemble of models. Particles consistent with what multiple models agree on get higher weight; outlier completions die off through resampling. Over time, the surviving particles represent sequences that all ensemble members collectively support — not a token-level average, but a sequence-level consensus.
The crucial difference from naive ensembling: reweighting happens at the sequence level, not independently at each token. A particle that has taken a path both models like accumulates higher weight than one that one model liked momentarily but the other didn’t. This lets the ensemble reason about full trajectories.
Concretely, if you have models p₁ and p₂, you might define the target distribution as proportional to p₁(x)^α · p₂(x)^(1-α) for some mixing weight α — a geometric mixture, also called a product of experts. SMC gives you a tractable way to approximately sample from this joint distribution, even when neither model can be inverted and the partition function is unknown.
What This Unlocks in Practice
The practical stakes are real. Practitioners already have access to dozens of open-weight models — Llama variants, Mistral, Qwen, Gemma — alongside proprietary APIs, each with different strengths. For a given task, model selection is often a coin flip based on benchmark proxy scores. Prompt sensitivity compounds the problem: the same task can see wildly different performance across paraphrased instructions.
SMC-based ensembling makes it possible to combine multiple models and multiple prompts in a single unified decoding pass. Rather than picking one model-prompt combination, you run particles through all of them simultaneously, letting the reweighting process determine which combinations produce sequences the ensemble agrees are good. This is particularly valuable when no single configuration dominates across all examples.
The framework also handles a subtlety that matters operationally: the models being ensembled don’t need to be the same size or even the same architecture. As long as each can score a partial sequence, it can participate as a reweighting signal. That means you can combine a large but slow model with a fast small model as a cheap proposal distribution — a form of speculative decoding with ensemble correction.
Challenges and What to Watch
SMC decoding isn’t free. Running N particles multiplies compute by roughly N, and resampling introduces its own overhead. The quality of the proposal distribution — how you sample candidate next tokens before reweighting — matters a lot; a bad proposal wastes particles on regions of sequence space nobody wants. Variance in weight estimates can cause particle collapse, where diversity collapses to a single hypothesis too early. Established SMC tricks like adaptive resampling schedules and twist functions that look ahead help here, and it’ll be worth watching how this paper handles those engineering choices.
There’s also the question of when this beats simpler alternatives. Minimum Bayes risk decoding, self-consistency voting, and best-of-N sampling are all competitive baselines that are considerably cheaper. The interesting regime for SMC ensembling is likely tasks where diversity across models or prompts is genuinely informative — not just any task, but ones where the ensemble components each have complementary failure modes.
As inference-time compute becomes the primary lever practitioners can pull (model weights are increasingly fixed), methods that do more with decoding — rather than training — are increasingly valuable. SMC sits at that intersection of principled inference and practical multi-model deployment, and the direction is worth following closely.