Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Queipo-de-Llano, Enrique, Arroyo, Álvaro, Barbero, Federico, Dong, Xiaowen, Bronstein, Michael, LeCun, Yann, Shwartz-Ziv, Ravid
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918329723125760
author Queipo-de-Llano, Enrique
Arroyo, Álvaro
Barbero, Federico
Dong, Xiaowen
Bronstein, Michael
LeCun, Yann
Shwartz-Ziv, Ravid
author_facet Queipo-de-Llano, Enrique
Arroyo, Álvaro
Barbero, Federico
Dong, Xiaowen
Bronstein, Michael
LeCun, Yann
Shwartz-Ziv, Ravid
contents Attention sinks and compression valleys have attracted significant attention as two puzzling phenomena in large language models, but have been studied in isolation. In this work, we present a surprising connection between attention sinks and compression valleys, tracing both to the formation of massive activations in the residual stream. We prove theoretically that massive activations necessarily produce representational compression and establish bounds on the resulting entropy reduction. Through experiments across several models (410M-120B parameters), we confirm that when the beginning-of-sequence token develops extreme activation norms in the middle layers, both compression valleys and attention sinks emerge simultaneously. Targeted ablation studies validate our theoretical predictions. This unified view motivates us to propose the Mix-Compress-Refine theory of information flow, as an attempt to explain how LLMs organize their computation in depth by controlling attention and representational compression via massive activations. Specifically, we posit that Transformer-based LLMs process tokens in three distinct phases: (1) broad mixing in the early layers, (2) compressed computation with limited mixing in the middle layers, and (3) selective refinement in the late layers. Our framework helps explain why embedding tasks perform best at intermediate layers, whereas generation tasks benefit from full-depth processing, clarifying differences in task-dependent representations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06477
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
Queipo-de-Llano, Enrique
Arroyo, Álvaro
Barbero, Federico
Dong, Xiaowen
Bronstein, Michael
LeCun, Yann
Shwartz-Ziv, Ravid
Machine Learning
Artificial Intelligence
Attention sinks and compression valleys have attracted significant attention as two puzzling phenomena in large language models, but have been studied in isolation. In this work, we present a surprising connection between attention sinks and compression valleys, tracing both to the formation of massive activations in the residual stream. We prove theoretically that massive activations necessarily produce representational compression and establish bounds on the resulting entropy reduction. Through experiments across several models (410M-120B parameters), we confirm that when the beginning-of-sequence token develops extreme activation norms in the middle layers, both compression valleys and attention sinks emerge simultaneously. Targeted ablation studies validate our theoretical predictions. This unified view motivates us to propose the Mix-Compress-Refine theory of information flow, as an attempt to explain how LLMs organize their computation in depth by controlling attention and representational compression via massive activations. Specifically, we posit that Transformer-based LLMs process tokens in three distinct phases: (1) broad mixing in the early layers, (2) compressed computation with limited mixing in the middle layers, and (3) selective refinement in the late layers. Our framework helps explain why embedding tasks perform best at intermediate layers, whereas generation tasks benefit from full-depth processing, clarifying differences in task-dependent representations.
title Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.06477