Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Zihan, Wang, Zekun, Zheng, Bo, Huang, Zeyu, Wen, Kaiyue, Yang, Songlin, Men, Rui, Yu, Le, Huang, Fei, Huang, Suozhi, Liu, Dayiheng, Zhou, Jingren, Lin, Junyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916731436400640
author Qiu, Zihan
Wang, Zekun
Zheng, Bo
Huang, Zeyu
Wen, Kaiyue
Yang, Songlin
Men, Rui
Yu, Le
Huang, Fei
Huang, Suozhi
Liu, Dayiheng
Zhou, Jingren
Lin, Junyang
author_facet Qiu, Zihan
Wang, Zekun
Zheng, Bo
Huang, Zeyu
Wen, Kaiyue
Yang, Songlin
Men, Rui
Yu, Le
Huang, Fei
Huang, Suozhi
Liu, Dayiheng
Zhou, Jingren
Lin, Junyang
contents Gating mechanisms have been widely utilized, from early models like LSTMs and Highway Networks to recent state space models, linear attention, and also softmax attention. Yet, existing literature rarely examines the specific effects of gating. In this work, we conduct comprehensive experiments to systematically investigate gating-augmented softmax attention variants. Specifically, we perform a comprehensive comparison over 30 variants of 15B Mixture-of-Experts (MoE) models and 1.7B dense models trained on a 3.5 trillion token dataset. Our central finding is that a simple modification-applying a head-specific sigmoid gate after the Scaled Dot-Product Attention (SDPA)-consistently improves performance. This modification also enhances training stability, tolerates larger learning rates, and improves scaling properties. By comparing various gating positions and computational variants, we attribute this effectiveness to two key factors: (1) introducing non-linearity upon the low-rank mapping in the softmax attention, and (2) applying query-dependent sparse gating scores to modulate the SDPA output. Notably, we find this sparse gating mechanism mitigates 'attention sink' and enhances long-context extrapolation performance, and we also release related $\href{https://github.com/qiuzh20/gated_attention}{codes}$ and $\href{https://huggingface.co/QwQZh/gated_attention}{models}$ to facilitate future research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06708
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Qiu, Zihan
Wang, Zekun
Zheng, Bo
Huang, Zeyu
Wen, Kaiyue
Yang, Songlin
Men, Rui
Yu, Le
Huang, Fei
Huang, Suozhi
Liu, Dayiheng
Zhou, Jingren
Lin, Junyang
Computation and Language
Gating mechanisms have been widely utilized, from early models like LSTMs and Highway Networks to recent state space models, linear attention, and also softmax attention. Yet, existing literature rarely examines the specific effects of gating. In this work, we conduct comprehensive experiments to systematically investigate gating-augmented softmax attention variants. Specifically, we perform a comprehensive comparison over 30 variants of 15B Mixture-of-Experts (MoE) models and 1.7B dense models trained on a 3.5 trillion token dataset. Our central finding is that a simple modification-applying a head-specific sigmoid gate after the Scaled Dot-Product Attention (SDPA)-consistently improves performance. This modification also enhances training stability, tolerates larger learning rates, and improves scaling properties. By comparing various gating positions and computational variants, we attribute this effectiveness to two key factors: (1) introducing non-linearity upon the low-rank mapping in the softmax attention, and (2) applying query-dependent sparse gating scores to modulate the SDPA output. Notably, we find this sparse gating mechanism mitigates 'attention sink' and enhances long-context extrapolation performance, and we also release related $\href{https://github.com/qiuzh20/gated_attention}{codes}$ and $\href{https://huggingface.co/QwQZh/gated_attention}{models}$ to facilitate future research.
title Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
topic Computation and Language
url https://arxiv.org/abs/2505.06708