Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
Fuente:
arXiv
Saved in:
| Main Authors: | Queipo-de-Llano, Enrique, Arroyo, Álvaro, Barbero, Federico, Dong, Xiaowen, Bronstein, Michael, LeCun, Yann, Shwartz-Ziv, Ravid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
by: Shani, Chen, et al.
Published: (2025)
by: Shani, Chen, et al.
Published: (2025)
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
by: Goldfeder, Judah, et al.
Published: (2026)
by: Goldfeder, Judah, et al.
Published: (2026)
The Entropy Enigma: Success and Failure of Entropy Minimization
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
by: Skean, Oscar, et al.
Published: (2024)
by: Skean, Oscar, et al.
Published: (2024)
On Training in Imagination
by: Timor, Nadav, et al.
Published: (2026)
by: Timor, Nadav, et al.
Published: (2026)
Variance-Covariance Regularization Improves Representation Learning
by: Zhu, Jiachen, et al.
Published: (2023)
by: Zhu, Jiachen, et al.
Published: (2023)
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
by: Sun, Shangwen, et al.
Published: (2026)
by: Sun, Shangwen, et al.
Published: (2026)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
by: Ioannides, Georgios, et al.
Published: (2025)
by: Ioannides, Georgios, et al.
Published: (2025)
An Information-Theoretic Perspective on Variance-Invariance-Covariance Regularization
by: Shwartz-Ziv, Ravid, et al.
Published: (2023)
by: Shwartz-Ziv, Ravid, et al.
Published: (2023)
Rate-In: Information-Driven Adaptive Dropout Rates for Improved Inference-Time Uncertainty Estimation
by: Zeevi, Tal, et al.
Published: (2024)
by: Zeevi, Tal, et al.
Published: (2024)
Learning to Compress: Local Rank and Information Compression in Deep Neural Networks
by: Patel, Niket, et al.
Published: (2024)
by: Patel, Niket, et al.
Published: (2024)
Just How Flexible are Neural Networks in Practice?
by: Shwartz-Ziv, Ravid, et al.
Published: (2024)
by: Shwartz-Ziv, Ravid, et al.
Published: (2024)
Layer by Layer: Uncovering Hidden Representations in Language Models
by: Skean, Oscar, et al.
Published: (2025)
by: Skean, Oscar, et al.
Published: (2025)
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
by: Arefin, Md Rifat, et al.
Published: (2024)
by: Arefin, Md Rifat, et al.
Published: (2024)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
by: Ioannides, Georgios, et al.
Published: (2026)
by: Ioannides, Georgios, et al.
Published: (2026)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
by: Sanyal, Sunny, et al.
Published: (2024)
by: Sanyal, Sunny, et al.
Published: (2024)
Fast and Exact Enumeration of Deep Networks Partitions Regions
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
by: Dawid, Anna, et al.
Published: (2023)
by: Dawid, Anna, et al.
Published: (2023)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
Two Sides of the Same Coin: Reaching Nontraditional Students
by: Berg, Steven L.
Published: (2005)
by: Berg, Steven L.
Published: (2005)
Distant and Distributed Learners Are Two Sides of the Same Coin.
by: Barron, Brette Barclay
Published: (2002)
by: Barron, Brette Barclay
Published: (2002)
Editorial: Metabolic Dysfunction and Alcohol—Two Sides of the Same Coin
by: Gustavo Ayares, et al.
Published: (2024)
by: Gustavo Ayares, et al.
Published: (2024)
Quantum Sensing and Quantum Error Correction: Two Sides of the Same Coin
by: Bao, Zhuoran, et al.
Published: (2026)
by: Bao, Zhuoran, et al.
Published: (2026)
Empowering and Resisting in a Sharing Economy: Two Sides of the Same Coin
by: Gicelda Julia Dal Bó
Published: (2019)
by: Gicelda Julia Dal Bó
Published: (2019)
Review Helpfulness Scores vs. Review Unhelpfulness Scores: Two Sides of the Same Coin or Different Coins?
by: Yu, Yinan, et al.
Published: (2024)
by: Yu, Yinan, et al.
Published: (2024)
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
Semantic Tube Prediction: Beating LLM Data Efficiency with JEPA
by: Huang, Hai, et al.
Published: (2026)
by: Huang, Hai, et al.
Published: (2026)
Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
by: Dupoux, Emmanuel, et al.
Published: (2026)
by: Dupoux, Emmanuel, et al.
Published: (2026)
Variance Covariance Regularization Enforces Pairwise Independence in Self-Supervised Representations
by: Mialon, Grégoire, et al.
Published: (2022)
by: Mialon, Grégoire, et al.
Published: (2022)
A hierarchical loss and its problems when classifying non-hierarchically
by: Wu, Cinna, et al.
Published: (2017)
by: Wu, Cinna, et al.
Published: (2017)
Editorial: Chronic Stimulation‐Induced Ataxia and Habituation after Thalamic DBS : Two Sides of the Same Coin?
by: A. Enrique Martinez‐Nunez, et al.
Published: (2025)
by: A. Enrique Martinez‐Nunez, et al.
Published: (2025)
Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
by: Wang, Yongjie, et al.
Published: (2025)
by: Wang, Yongjie, et al.
Published: (2025)
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
by: LeVi, Amit, et al.
Published: (2025)
by: LeVi, Amit, et al.
Published: (2025)
LEAF: Unveiling Two Sides of the Same Coin in Semi-supervised Facial Expression Recognition
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
Catalyst: Instructional Development and Teaching Library Media Skills: Two Sides of the Same Coin?
Published: (1986)
Published: (1986)
School Library Media Specialist/School Board Member: Two Sides of the Same Coin.
by: Nutt, Pam
Published: (2003)
by: Nutt, Pam
Published: (2003)
Two Sides of the Same Coin‐Baclofen‐Induced Intoxication in a Chronic Hemodialysis Patient
by: Gerry George Mathew, et al.
Published: (2025)
by: Gerry George Mathew, et al.
Published: (2025)
Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models
by: Paech, Samuel, et al.
Published: (2025)
by: Paech, Samuel, et al.
Published: (2025)
Similar Items
-
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024) -
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
by: Shani, Chen, et al.
Published: (2025) -
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
by: Goldfeder, Judah, et al.
Published: (2026) -
The Entropy Enigma: Success and Failure of Entropy Minimization
by: Press, Ori, et al.
Published: (2024) -
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
by: Skean, Oscar, et al.
Published: (2024)