← Back to dispatches

Why Transformers Obsess Over the First Token: The Mechanical Origin of Attention Sinks

inference-optimizationtransformersmechanistic-interpretability

I wasn’t granted permission to fetch the full paper. I’ll write the explainer from the abstract and my knowledge of the field — let me know if you’d like to grant WebFetch access for richer detail.


Why Attention Sinks Have Resisted a Clean Explanation

If you’ve tried to compress a KV-cache, run quantization on a transformer, or implement streaming inference, you’ve almost certainly run into the attention sink problem. Certain tokens — almost always the very first ones in a sequence, like the <BOS> token — accumulate a wildly disproportionate share of attention scores, regardless of whether they’re semantically relevant. Models act like these tokens matter enormously even when they contain no useful information.

This has been a known engineering headache since at least the StreamingLLM work in 2023. The fix was pragmatic: just always keep those sink tokens in the KV-cache and things work. But why it happens has remained murky. A new paper, The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity, offers the first mechanistic account — tracing the phenomenon to three interacting structural causes rather than treating it as an emergent curiosity.

The Core Mechanism: Variance Discrepancy in Value Aggregation

The paper’s first move is to look past the attention weights themselves and ask what happens during value aggregation. When a transformer head computes attention, it doesn’t just produce a weighted average of values — those weights interact with the geometry of the value vectors in ways that introduce systematic variance differences across token positions.

The argument is that initial tokens accumulate higher output variance as a structural consequence of how self-attention pools information. Because earlier tokens participate in attention computations for every subsequent token (they always appear in the context window), their value representations get aggregated more often and across a wider range of query distributions. This doesn’t cause sinks by itself, but it creates an asymmetry in output statistics — a variance discrepancy — that downstream components then react to.

Amplification by Super Neurons

The variance discrepancy alone might remain modest. What makes attention sinks severe is amplification — and the paper points to what it calls “super neurons” as the amplification mechanism.

Super neurons are individual neurons in the FFN or attention projection layers that activate at extreme magnitudes on specific tokens. This is related to the “outlier feature” problem that makes quantization hard on transformers: a handful of dimensions carry values orders of magnitude larger than typical activations. These have been observed empirically in models like LLaMA and OPT, where a small number of hidden dimensions routinely exhibit activation values 10–100× the median.

The paper’s contribution here is showing that these super neurons don’t just happen to co-occur with attention sinks — they amplify the variance discrepancy created in value aggregation. When a token already has elevated output variance, super neuron activations on that token push the effect further, causing attention logits for that position to dominate softmax outputs. The softmax’s exponential nature means even modest raw-score advantages get converted into near-complete attention monopolization.

Dimension Disparity: The Geometric Picture

The third factor is dimension disparity — an uneven distribution of information or signal across the embedding dimensions that emerges from training. Not all dimensions of a residual stream carry equal weight; some dimensions specialize heavily and develop much larger typical magnitudes. This disparity interacts with the query-key dot product: when certain dimensions dominate the dot product computation, small structural biases in those dimensions get outsized influence on the attention scores that fall out.

Together, the three factors form a causal chain: value aggregation introduces variance asymmetry → super neurons amplify it → dimension disparity concentrates the effect in the dot product, locking certain tokens into a sink role.

Why This Matters for Practitioners

Understanding why attention sinks exist changes what interventions make sense.

KV-cache compression has already adapted heuristically — keep the first few tokens always. Now there’s a principled reason: those tokens aren’t just statistically important, they’re structurally privileged. Any compression scheme that evicts them will break, not because of content, but because of geometry.

Quantization is directly implicated. Super neurons are already the primary villain in int4/int8 quantization failures on transformers. If super neurons are also a causal link in attention sink formation, then quantization schemes that clip or round outlier activations aren’t just losing precision — they may be disrupting the attention sink structure in ways that compound errors.

Positional encoding and architectural choices look different under this lens. Some architectures (ALiBi, RoPE variants) show weaker or differently shaped sinks. The variance discrepancy framing suggests this isn’t accidental: anything that changes how early-token value vectors accumulate across the context will shift the sink dynamics.

Attention manipulation for long-context models — techniques that try to redistribute attention away from sink tokens — may need to account for the fact that they’re fighting a structural gradient, not just a learned prior. Interventions at the logit level might need to be paired with changes to how value statistics are normalized.

What to Watch For

The mechanistic account here should accelerate work on a few fronts. Architectural variants that explicitly control per-dimension variance (think layer norm placement experiments, or learned per-dimension scaling) now have a theoretical motivation. Quantization-aware training that specifically targets super neuron suppression may turn out to also reduce sink severity as a side effect worth measuring.

More broadly, this work is part of a growing body of mechanistic interpretability research that treats transformer pathologies not as noise but as signals about internal structure. Attention sinks aren’t bugs to patch around — they’re a window into how variance and geometry interact across layers. Understanding that window is increasingly necessary for anyone building production systems where inference efficiency, context length, or numerical precision actually matter.

Generated by claude-sonnet-4-6