SIMD All the Way Down: Running LLMs on Your CPU with Ternary Weights
I wasn’t granted access to fetch the paper. I’ll write the explainer based on the abstract and my knowledge of this technical domain.
Why Running LLMs Locally Is Harder Than It Should Be
Your laptop has more raw compute than the machines that trained early neural networks. Yet running a capable large language model locally still typically requires either a high-end GPU, a cloud subscription, or patience measured in minutes per token. The bottleneck isn’t fundamentally about compute — it’s about memory bandwidth and the mismatch between how modern models are represented and what consumer CPUs are actually good at.
Ternary neural networks attack this mismatch at the source. By constraining every model weight to one of three values — {-1, 0, +1} — they promise to turn expensive floating-point multiply-accumulate operations into cheap additions, subtractions, and skips. The theoretical efficiency gain is enormous. The practical efficiency gain, until recently, has been much smaller, because mainstream inference frameworks weren’t built to exploit the structure. Litespark is a focused attempt to close that gap with hand-tuned SIMD kernels targeting the CPUs already sitting in over a billion personal computers.
The Ternary Opportunity
When a weight is 0, you skip the accumulation entirely. When it’s +1 or -1, you add or subtract the activation — no multiplication needed. This sounds simple, but realizing it in practice requires solving a packing problem: how do you store and process weights that only need 1.58 bits each using 8-, 16-, or 32-bit SIMD lanes?
The standard approach in frameworks like llama.cpp is to quantize aggressively, but generic quantization assumes arbitrary weight values and builds in the infrastructure for dequantization at runtime. Ternary weights break that assumption in a useful direction — you don’t need to recover a float, you need to branch on three discrete states — but most frameworks treat them as a special case of 2-bit quantization rather than a structurally distinct representation. The result is that the theoretical multiply-free arithmetic either never materializes, or materializes only partially after expensive setup.
Litespark’s core contribution is a set of SIMD kernels — targeting x86 extensions like AVX2 and AVX-512 — that are written specifically for ternary arithmetic from the ground up, rather than adapted from float or generic-quantization paths.
How the Kernels Work
The key insight is weight packing. Ternary weights can be encoded in two bits (or, with a sign-and-zero scheme, in two parallel bitmasks: one for sign, one for presence). By packing multiple weights into a single SIMD register and using bitwise operations to dispatch additions and subtractions in bulk, you can process many weights per clock cycle without ever entering floating-point territory for the core multiply-accumulate loop.
AVX2 gives you 256-bit registers, meaning you can pack and process 128 ternary weights simultaneously per register operation if your encoding is tight. AVX-512, available on recent Intel and AMD desktop CPUs, doubles that to 512-bit registers. The kernels exploit this by organizing matrix-vector products — the dominant operation in transformer inference — as sequences of bitwise ANDs, population counts, and integer additions rather than fmul/fmadd chains.
The zero weight is especially valuable: roughly half of ternary model weights are zero in well-trained models (a property that emerges from training, not just quantization). A kernel aware of this structure can skip entire chunks of work that a generic 2-bit kernel would still process.
What This Means for Real Hardware
Consumer CPUs have lagged GPUs for LLM inference primarily because of memory bandwidth, not FLOPs. A ternary model is dramatically smaller in memory than a float16 equivalent — a 7B parameter model that would occupy ~14 GB in float16 shrinks to roughly 1.7 GB in packed ternary format. That fits comfortably in L3 cache or system RAM with fast sequential access patterns, which is exactly what SIMD-friendly matrix kernels need.
The implication is that the inference bottleneck shifts. Rather than waiting on GPU VRAM bandwidth or paying for API calls, a ternary model with proper CPU kernels can leverage the memory hierarchy of an ordinary desktop or laptop. Litespark targets this regime directly, making the case that the “billion underutilized personal computers” in the abstract aren’t underutilized because CPUs are slow — they’re underutilized because no one wrote the right kernels.
What to Watch For
The broader ecosystem is moving in this direction. Models like BitNet b1.58 from Microsoft Research demonstrated that ternary-weight transformers can reach competitive quality, giving the inference optimization work a viable model family to target. The missing piece has been the runtime layer, and projects like Litespark are filling it.
For developers, the practical implication is that local inference of capable models on CPU-only hardware is becoming less exotic. If ternary kernels continue maturing — particularly with AVX-512 coverage expanding as Zen 4 and Intel’s recent microarchitectures become mainstream — the calculus on when to reach for a GPU or cloud API shifts. Embedded deployment, air-gapped environments, and privacy-sensitive applications all benefit from inference that runs well on a stock developer machine.
The deeper lesson is one of abstraction mismatch: powerful hardware capabilities go unrealized when the software stack wasn’t designed with the underlying structure in mind. Ternary weights aren’t just quantized weights — they’re a different arithmetic regime, and they need kernels that treat them as such.