Saved in:
| Main Authors: | Li, Siquan, Jiang, Kaiqi, Sun, Jiacheng, Hu, Tianyang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.06611 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attention Sinks and Outliers in Attention Residuals
by: Luo, Haozheng, et al.
Published: (2026)
by: Luo, Haozheng, et al.
Published: (2026)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026)
by: Chen, Yihong, et al.
Published: (2026)
On the Existence and Behavior of Secondary Attention Sinks
by: Wong, Jeffrey T. H., et al.
Published: (2025)
by: Wong, Jeffrey T. H., et al.
Published: (2025)
Transformers Are Born Biased: Structural Inductive Biases at Random Initialization and Their Practical Consequences
by: Li, Siquan, et al.
Published: (2026)
by: Li, Siquan, et al.
Published: (2026)
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
On the Role of Preference Variance in Preference Optimization
by: Guo, Jiacheng, et al.
Published: (2025)
by: Guo, Jiacheng, et al.
Published: (2025)
When Attention Sink Emerges in Language Models: An Empirical View
by: Gu, Xiangming, et al.
Published: (2024)
by: Gu, Xiangming, et al.
Published: (2024)
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
by: Zhang, Stephen, et al.
Published: (2025)
by: Zhang, Stephen, et al.
Published: (2025)
A Minimum Variance Path Principle for Accurate and Stable Score-Based Density Ratio Estimation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
Neuronal Attention Circuit (NAC) for Representation Learning
by: Razzaq, Waleed, et al.
Published: (2025)
by: Razzaq, Waleed, et al.
Published: (2025)
Streaming Attention Approximation via Discrepancy Theory
by: Kochetkova, Ekaterina, et al.
Published: (2025)
by: Kochetkova, Ekaterina, et al.
Published: (2025)
When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models
by: Liu, Junnan, et al.
Published: (2026)
by: Liu, Junnan, et al.
Published: (2026)
ProofAug: Efficient Neural Theorem Proving via Fine-grained Proof Structure Analysis
by: Liu, Haoxiong, et al.
Published: (2025)
by: Liu, Haoxiong, et al.
Published: (2025)
Neuronal Stochastic Attention Circuit (NSAC) for Probabilistic Representation Learning
by: Razzaq, Waleed, et al.
Published: (2026)
by: Razzaq, Waleed, et al.
Published: (2026)
LATTLE: LLM Attention Transplant for Transfer Learning of Tabular Data Across Disparate Domains
by: Kowsar, Ibna, et al.
Published: (2025)
by: Kowsar, Ibna, et al.
Published: (2025)
Federated Unlearning in the Wild: Rethinking Fairness and Data Discrepancy
by: Huang, ZiHeng, et al.
Published: (2025)
by: Huang, ZiHeng, et al.
Published: (2025)
Strassen Attention, Split VC Dimension and Compositionality in Transformers
by: Kozachinskiy, Alexander, et al.
Published: (2025)
by: Kozachinskiy, Alexander, et al.
Published: (2025)
Graph Dimension Attention Networks for Enterprise Credit Assessment
by: Wei, Shaopeng, et al.
Published: (2024)
by: Wei, Shaopeng, et al.
Published: (2024)
Self-Supervised Representation Learning for Geospatial Objects: A Survey
by: Chen, Yile, et al.
Published: (2024)
by: Chen, Yile, et al.
Published: (2024)
Decomposing Attention To Find Context-Sensitive Neurons
by: Gibson, Alex
Published: (2025)
by: Gibson, Alex
Published: (2025)
GCFX: Generative Counterfactual Explanations for Deep Graph Models at the Model Level
by: Hu, Jinlong, et al.
Published: (2026)
by: Hu, Jinlong, et al.
Published: (2026)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
by: Luo, Yu, et al.
Published: (2026)
by: Luo, Yu, et al.
Published: (2026)
Memorization Sinks: Isolating Memorization during LLM Training
by: Ghosal, Gaurav R., et al.
Published: (2025)
by: Ghosal, Gaurav R., et al.
Published: (2025)
Detection vs. Execution: Single-Bucket Probes Miss Half the Mamba-2 State Sink
by: Jiang, Yuhang
Published: (2026)
by: Jiang, Yuhang
Published: (2026)
Sink-Aware Pruning for Diffusion Language Models
by: Myrzakhan, Aidar, et al.
Published: (2026)
by: Myrzakhan, Aidar, et al.
Published: (2026)
Dual Randomized Smoothing: Beyond Global Noise Variance
by: Sun, Chenhao, et al.
Published: (2025)
by: Sun, Chenhao, et al.
Published: (2025)
Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
by: Wang, Zhiwei, et al.
Published: (2024)
by: Wang, Zhiwei, et al.
Published: (2024)
Taking Shortcuts for Categorical VQA Using Super Neurons
by: Musacchio, Pierre, et al.
Published: (2026)
by: Musacchio, Pierre, et al.
Published: (2026)
Unveiling and Controlling Anomalous Attention Distribution in Transformers
by: Yan, Ruiqing, et al.
Published: (2024)
by: Yan, Ruiqing, et al.
Published: (2024)
A Variance-Reduced Cubic-Regularized Newton for Policy Optimization
by: Sun, Cheng, et al.
Published: (2025)
by: Sun, Cheng, et al.
Published: (2025)
Incorporating Domain Differential Equations into Graph Convolutional Networks to Lower Generalization Discrepancy
by: Sun, Yue, et al.
Published: (2024)
by: Sun, Yue, et al.
Published: (2024)
On Vanishing Variance in Transformer Length Generalization
by: Li, Ruining, et al.
Published: (2025)
by: Li, Ruining, et al.
Published: (2025)
Fairness in Survival Analysis: A Novel Conditional Mutual Information Augmentation Approach
by: Xie, Tianyang, et al.
Published: (2025)
by: Xie, Tianyang, et al.
Published: (2025)
Magnitude-based Neuron Pruning for Backdoor Defens
by: Li, Nan, et al.
Published: (2024)
by: Li, Nan, et al.
Published: (2024)
Weighted Graph Structure Learning with Attention Denoising for Node Classification
by: Wang, Tingting, et al.
Published: (2025)
by: Wang, Tingting, et al.
Published: (2025)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
by: Huang, Luke J., et al.
Published: (2026)
by: Huang, Luke J., et al.
Published: (2026)
AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies
by: Hu, Xixi, et al.
Published: (2024)
by: Hu, Xixi, et al.
Published: (2024)
Bias Detection via Maximum Subgroup Discrepancy
by: Němeček, Jiří, et al.
Published: (2025)
by: Němeček, Jiří, et al.
Published: (2025)
Discrepancy-Aware Graph Mask Auto-Encoder
by: Zheng, Ziyu, et al.
Published: (2025)
by: Zheng, Ziyu, et al.
Published: (2025)
Similar Items
-
Attention Sinks and Outliers in Attention Residuals
by: Luo, Haozheng, et al.
Published: (2026) -
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026) -
On the Existence and Behavior of Secondary Attention Sinks
by: Wong, Jeffrey T. H., et al.
Published: (2025) -
Transformers Are Born Biased: Structural Inductive Biases at Random Initialization and Their Practical Consequences
by: Li, Siquan, et al.
Published: (2026) -
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)