A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
Fuente:
arXiv
Saved in:
| Main Authors: | Ran-Milo, Yuval, Ofek, Hila, Mendel, Shahar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
by: Ran-Milo, Yuval
Published: (2026)
by: Ran-Milo, Yuval
Published: (2026)
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
by: Ran-Milo, Yuval, et al.
Published: (2026)
by: Ran-Milo, Yuval, et al.
Published: (2026)
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
Do Neural Networks Need Gradient Descent to Generalize? A Theoretical Study
by: Alexander, Yotam, et al.
Published: (2025)
by: Alexander, Yotam, et al.
Published: (2025)
Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation
by: Mutisya, Hillary, et al.
Published: (2026)
by: Mutisya, Hillary, et al.
Published: (2026)
How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability
by: García-Carrasco, Jorge, et al.
Published: (2024)
by: García-Carrasco, Jorge, et al.
Published: (2024)
Attention Sinks and Outliers in Attention Residuals
by: Luo, Haozheng, et al.
Published: (2026)
by: Luo, Haozheng, et al.
Published: (2026)
ASAP: Attention Sink Anchored Pruning
by: Lee, Jaehyuk, et al.
Published: (2026)
by: Lee, Jaehyuk, et al.
Published: (2026)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
by: Son, Seungwoo, et al.
Published: (2024)
by: Son, Seungwoo, et al.
Published: (2024)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
by: Braun, Dan, et al.
Published: (2025)
by: Braun, Dan, et al.
Published: (2025)
Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
by: He, Zhengfu, et al.
Published: (2024)
by: He, Zhengfu, et al.
Published: (2024)
A Broader View of Thompson Sampling
by: Qu, Yanlin, et al.
Published: (2025)
by: Qu, Yanlin, et al.
Published: (2025)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026)
by: Chen, Yihong, et al.
Published: (2026)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Stochastic Parroting in Temporal Attention -- Regulating the Diagonal Sink
by: Hankemeier, Victoria, et al.
Published: (2026)
by: Hankemeier, Victoria, et al.
Published: (2026)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
by: Zuhri, Zayd M. K., et al.
Published: (2025)
by: Zuhri, Zayd M. K., et al.
Published: (2025)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
by: Bushnaq, Lucius, et al.
Published: (2024)
by: Bushnaq, Lucius, et al.
Published: (2024)
On the Existence and Behavior of Secondary Attention Sinks
by: Wong, Jeffrey T. H., et al.
Published: (2025)
by: Wong, Jeffrey T. H., et al.
Published: (2025)
Mamba Knockout for Unraveling Factual Information Flow
by: Endy, Nir, et al.
Published: (2025)
by: Endy, Nir, et al.
Published: (2025)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025)
by: Baroni, Luca, et al.
Published: (2025)
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
by: Bali, Karan, et al.
Published: (2026)
by: Bali, Karan, et al.
Published: (2026)
Mechanistic Analysis of Circuit Preservation in Federated Learning
by: Haseeb, Muhammad, et al.
Published: (2025)
by: Haseeb, Muhammad, et al.
Published: (2025)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
by: Shang, Bingqi, et al.
Published: (2025)
by: Shang, Bingqi, et al.
Published: (2025)
On Mechanistic Circuits for Extractive Question-Answering
by: Basu, Samyadeep, et al.
Published: (2025)
by: Basu, Samyadeep, et al.
Published: (2025)
Spectral Path Regression: Directional Chebyshev Harmonics for Interpretable Tabular Learning
by: Coombs, Milo
Published: (2026)
by: Coombs, Milo
Published: (2026)
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
by: Zhang, Stephen, et al.
Published: (2025)
by: Zhang, Stephen, et al.
Published: (2025)
From Posterior Sampling to Meaningful Diversity in Image Restoration
by: Cohen, Noa, et al.
Published: (2023)
by: Cohen, Noa, et al.
Published: (2023)
Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
by: Haskins, Reilly, et al.
Published: (2025)
by: Haskins, Reilly, et al.
Published: (2025)
Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
by: Huang, Xingyue, et al.
Published: (2026)
by: Huang, Xingyue, et al.
Published: (2026)
A Generative Approach for Semantic Auditing of Electronic Health Records
by: Girshovitz, Irena, et al.
Published: (2025)
by: Girshovitz, Irena, et al.
Published: (2025)
Physics-informed Blind Reconstruction of Dense Fields from Sparse Measurements using Neural Networks with a Differentiable Simulator
by: Aloni, Ofek, et al.
Published: (2026)
by: Aloni, Ofek, et al.
Published: (2026)
Provable Benefits of Complex Parameterizations for Structured State Space Models
by: Ran-Milo, Yuval, et al.
Published: (2024)
by: Ran-Milo, Yuval, et al.
Published: (2024)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
by: Binkowski, Jakub, et al.
Published: (2026)
by: Binkowski, Jakub, et al.
Published: (2026)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
Identifying a Circuit for Verb Conjugation in GPT-2
by: Africa, David Demitri
Published: (2025)
by: Africa, David Demitri
Published: (2025)
When Attention Sink Emerges in Language Models: An Empirical View
by: Gu, Xiangming, et al.
Published: (2024)
by: Gu, Xiangming, et al.
Published: (2024)
The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity
by: Li, Siquan, et al.
Published: (2026)
by: Li, Siquan, et al.
Published: (2026)
Boosting Few-Pixel Robustness Verification via Covering Verification Designs
by: Shapira, Yuval, et al.
Published: (2024)
by: Shapira, Yuval, et al.
Published: (2024)
Similar Items
-
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
by: Ran-Milo, Yuval
Published: (2026) -
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
by: Ran-Milo, Yuval, et al.
Published: (2026) -
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
by: Su, Zunhai, et al.
Published: (2026) -
Do Neural Networks Need Gradient Descent to Generalize? A Theoretical Study
by: Alexander, Yotam, et al.
Published: (2025) -
Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation
by: Mutisya, Hillary, et al.
Published: (2026)