Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Zihan, Wang, Zekun, Zheng, Bo, Huang, Zeyu, Wen, Kaiyue, Yang, Songlin, Men, Rui, Yu, Le, Huang, Fei, Huang, Suozhi, Liu, Dayiheng, Zhou, Jingren, Lin, Junyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
by: Qiu, Zihan, et al.
Published: (2025)
by: Qiu, Zihan, et al.
Published: (2025)
A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Transformer Training
by: Qiu, Zihan, et al.
Published: (2026)
by: Qiu, Zihan, et al.
Published: (2026)
Attention Sinks and Outliers in Attention Residuals
by: Luo, Haozheng, et al.
Published: (2026)
by: Luo, Haozheng, et al.
Published: (2026)
Scaling Attention via Feature Sparsity
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
by: Huang, Xingyue, et al.
Published: (2026)
by: Huang, Xingyue, et al.
Published: (2026)
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training
by: Tang, Shengkun, et al.
Published: (2026)
by: Tang, Shengkun, et al.
Published: (2026)
Unlocking Emergent Modularity in Large Language Models
by: Qiu, Zihan, et al.
Published: (2023)
by: Qiu, Zihan, et al.
Published: (2023)
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
by: Sun, Shangwen, et al.
Published: (2026)
by: Sun, Shangwen, et al.
Published: (2026)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention Sink
by: Liu, Guozhi, et al.
Published: (2026)
by: Liu, Guozhi, et al.
Published: (2026)
SinkTrack: Attention Sink based Context Anchoring for Large Language Models
by: Liu, Xu, et al.
Published: (2026)
by: Liu, Xu, et al.
Published: (2026)
Language Confusion Gate: Language-Aware Decoding Through Model Self-Distillation
by: Zhang, Collin, et al.
Published: (2025)
by: Zhang, Collin, et al.
Published: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning
by: Sok, Jaewon, et al.
Published: (2026)
by: Sok, Jaewon, et al.
Published: (2026)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
Empirical Study on Updating Key-Value Memories in Transformer Feed-forward Layers
by: Qiu, Zihan, et al.
Published: (2024)
by: Qiu, Zihan, et al.
Published: (2024)
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
by: Choi, Jiho, et al.
Published: (2026)
by: Choi, Jiho, et al.
Published: (2026)
HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning
by: Li, Xiaoyuan, et al.
Published: (2025)
by: Li, Xiaoyuan, et al.
Published: (2025)
Attention Sinks in Diffusion Language Models
by: Rulli, Maximo Eduardo, et al.
Published: (2025)
by: Rulli, Maximo Eduardo, et al.
Published: (2025)
ASAP: Attention Sink Anchored Pruning
by: Lee, Jaehyuk, et al.
Published: (2026)
by: Lee, Jaehyuk, et al.
Published: (2026)
On the Existence and Behavior of Secondary Attention Sinks
by: Wong, Jeffrey T. H., et al.
Published: (2025)
by: Wong, Jeffrey T. H., et al.
Published: (2025)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
ActFormer: Scalable Collaborative Perception via Active Queries
by: Huang, Suozhi, et al.
Published: (2024)
by: Huang, Suozhi, et al.
Published: (2024)
GateAttentionPose: Enhancing Pose Estimation with Agent Attention and Improved Gated Convolutions
by: Feng, Liang, et al.
Published: (2024)
by: Feng, Liang, et al.
Published: (2024)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
by: Binkowski, Jakub, et al.
Published: (2026)
by: Binkowski, Jakub, et al.
Published: (2026)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026)
by: Chen, Yihong, et al.
Published: (2026)
CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction
by: Li, Zixuan, et al.
Published: (2025)
by: Li, Zixuan, et al.
Published: (2025)
Efficient Streaming Language Models with Attention Sinks
by: Xiao, Guangxuan, et al.
Published: (2023)
by: Xiao, Guangxuan, et al.
Published: (2023)
Spectral Filters, Dark Signals, and Attention Sinks
by: Cancedda, Nicola
Published: (2024)
by: Cancedda, Nicola
Published: (2024)
Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering
by: Luo, Jiayi, et al.
Published: (2026)
by: Luo, Jiayi, et al.
Published: (2026)
Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-based Optimizers
by: Wang, Liangyu, et al.
Published: (2026)
by: Wang, Liangyu, et al.
Published: (2026)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
by: Yu, Zhongzhi, et al.
Published: (2024)
by: Yu, Zhongzhi, et al.
Published: (2024)
Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning
by: Yi, Shuai, et al.
Published: (2026)
by: Yi, Shuai, et al.
Published: (2026)
SoLA: Leveraging Soft Activation Sparsity and Low-Rank Decomposition for Large Language Model Compression
by: Huang, Xinhao, et al.
Published: (2026)
by: Huang, Xinhao, et al.
Published: (2026)
Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
by: Zhang, Yanzhao, et al.
Published: (2025)
by: Zhang, Yanzhao, et al.
Published: (2025)
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
Power-based Partial Attention: Bridging Linear-Complexity and Full Attention
by: Huang, Yufeng
Published: (2026)
by: Huang, Yufeng
Published: (2026)
Similar Items
-
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
by: Qiu, Zihan, et al.
Published: (2025) -
A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Transformer Training
by: Qiu, Zihan, et al.
Published: (2026) -
Attention Sinks and Outliers in Attention Residuals
by: Luo, Haozheng, et al.
Published: (2026) -
Scaling Attention via Feature Sparsity
by: Xie, Yan, et al.
Published: (2026) -
Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
by: Huang, Xingyue, et al.
Published: (2026)