Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention
Fuente:
arXiv
Saved in:
| Main Authors: | Súkeník, Peter, Amado, Cristina López, Lampert, Christoph H., Mondelli, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers
by: Súkeník, Peter, et al.
Published: (2025)
by: Súkeník, Peter, et al.
Published: (2025)
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
by: Súkeník, Peter, et al.
Published: (2024)
by: Súkeník, Peter, et al.
Published: (2024)
Average gradient outer product as a mechanism for deep neural collapse
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Intriguing Properties of Input-dependent Randomized Smoothing
by: Súkeník, Peter, et al.
Published: (2021)
by: Súkeník, Peter, et al.
Published: (2021)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
by: Jacot, Arthur, et al.
Published: (2024)
by: Jacot, Arthur, et al.
Published: (2024)
Learning Quantized Continuous Controllers for Integer Hardware
by: Kresse, Fabian, et al.
Published: (2025)
by: Kresse, Fabian, et al.
Published: (2025)
Tucker Attention: A generalization of approximate attention mechanisms
by: Klein, Timon, et al.
Published: (2026)
by: Klein, Timon, et al.
Published: (2026)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026)
by: Chen, Yihong, et al.
Published: (2026)
SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models
by: Liu, Junnan, et al.
Published: (2026)
by: Liu, Junnan, et al.
Published: (2026)
Attention Sinks and Outliers in Attention Residuals
by: Luo, Haozheng, et al.
Published: (2026)
by: Luo, Haozheng, et al.
Published: (2026)
Detection vs. Execution: Single-Bucket Probes Miss Half the Mamba-2 State Sink
by: Jiang, Yuhang
Published: (2026)
by: Jiang, Yuhang
Published: (2026)
On the Existence and Behavior of Secondary Attention Sinks
by: Wong, Jeffrey T. H., et al.
Published: (2025)
by: Wong, Jeffrey T. H., et al.
Published: (2025)
Memorization Sinks: Isolating Memorization during LLM Training
by: Ghosal, Gaurav R., et al.
Published: (2025)
by: Ghosal, Gaurav R., et al.
Published: (2025)
Reorganizing attention-space geometry with expressive attention
by: Gros, Claudius
Published: (2024)
by: Gros, Claudius
Published: (2024)
FLUID: Continuous-Time Hyperconnected Sparse Transformer for Sink-Free Learning
by: Razzaq, Waleed, et al.
Published: (2026)
by: Razzaq, Waleed, et al.
Published: (2026)
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
Poly-attention: a general scheme for higher-order self-attention
by: Chakrabarti, Sayak, et al.
Published: (2026)
by: Chakrabarti, Sayak, et al.
Published: (2026)
PRISM: Lightweight Multivariate Time-Series Classification through Symmetric Multi-Resolution Convolutional Layers
by: Zucchi, Federico, et al.
Published: (2025)
by: Zucchi, Federico, et al.
Published: (2025)
The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity
by: Li, Siquan, et al.
Published: (2026)
by: Li, Siquan, et al.
Published: (2026)
Supervised learning pays attention
by: Craig, Erin, et al.
Published: (2025)
by: Craig, Erin, et al.
Published: (2025)
Sink-Aware Pruning for Diffusion Language Models
by: Myrzakhan, Aidar, et al.
Published: (2026)
by: Myrzakhan, Aidar, et al.
Published: (2026)
Enhancing short-term traffic prediction by integrating trends and fluctuations with attention mechanism
by: Das, Adway, et al.
Published: (2025)
by: Das, Adway, et al.
Published: (2025)
Small transformer architectures for task switching
by: Gros, Claudius
Published: (2025)
by: Gros, Claudius
Published: (2025)
When Attention Sink Emerges in Language Models: An Empirical View
by: Gu, Xiangming, et al.
Published: (2024)
by: Gu, Xiangming, et al.
Published: (2024)
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
by: Zhang, Stephen, et al.
Published: (2025)
by: Zhang, Stephen, et al.
Published: (2025)
EEG motor imagery decoding: A framework for comparative analysis with channel attention mechanisms
by: Wimpff, Martin, et al.
Published: (2023)
by: Wimpff, Martin, et al.
Published: (2023)
An end-to-end attention-based approach for learning on graphs
by: Buterez, David, et al.
Published: (2024)
by: Buterez, David, et al.
Published: (2024)
Short window attention enables long-term memorization
by: Cabannes, Loïc, et al.
Published: (2025)
by: Cabannes, Loïc, et al.
Published: (2025)
Cross-attentive Cohesive Subgraph Embedding to Mitigate Oversquashing in GNNs
by: Hossain, Tanvir, et al.
Published: (2026)
by: Hossain, Tanvir, et al.
Published: (2026)
Goal Recognition as Reinforcement Learning
by: Amado, Leonardo Rosa, et al.
Published: (2022)
by: Amado, Leonardo Rosa, et al.
Published: (2022)
When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
GQA-μP: The maximal parameterization update for grouped query attention
by: Chickering, Kyle R., et al.
Published: (2026)
by: Chickering, Kyle R., et al.
Published: (2026)
GRC-Net: Gram Residual Co-attention Net for epilepsy prediction
by: You, Bihao, et al.
Published: (2025)
by: You, Bihao, et al.
Published: (2025)
Static and multivariate-temporal attentive fusion transformer for readmission risk prediction
by: Sun, Zhe, et al.
Published: (2024)
by: Sun, Zhe, et al.
Published: (2024)
Predicting the Geolocation of Tweets Using transformer models on Customized Data
by: Lutsai, Kateryna, et al.
Published: (2023)
by: Lutsai, Kateryna, et al.
Published: (2023)
Optimization of bi-directional gated loop cell based on multi-head attention mechanism for SSD health state classification model
by: Wen, Zhizhao, et al.
Published: (2025)
by: Wen, Zhizhao, et al.
Published: (2025)
Fairness vs Performance: Characterizing the Pareto Frontier of Algorithmic Decision Systems
by: Wilms, Mieke, et al.
Published: (2026)
by: Wilms, Mieke, et al.
Published: (2026)
HeatGen: A Guided Diffusion Framework for Multiphysics Heat Sink Design Optimization
by: Keramati, Hadi, et al.
Published: (2025)
by: Keramati, Hadi, et al.
Published: (2025)
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
by: Shin, Seungjun, et al.
Published: (2025)
by: Shin, Seungjun, et al.
Published: (2025)
Multi-layer Cross-attention is Provably Optimal for Multi-modal In-context Learning
by: Barnfield, Nicholas, et al.
Published: (2026)
by: Barnfield, Nicholas, et al.
Published: (2026)
Similar Items
-
Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers
by: Súkeník, Peter, et al.
Published: (2025) -
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
by: Súkeník, Peter, et al.
Published: (2024) -
Average gradient outer product as a mechanism for deep neural collapse
by: Beaglehole, Daniel, et al.
Published: (2024) -
Intriguing Properties of Input-dependent Randomized Smoothing
by: Súkeník, Peter, et al.
Published: (2021) -
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
by: Jacot, Arthur, et al.
Published: (2024)