Grouped Differential Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lim, Junghwan, Lee, Sungmin, Kim, Dongseok, Cheung, Wai Ting, Kim, Beomgyu, Kim, Taehwan, Lee, Haesol, Lee, Junhyeok, Oh, Dongpin, Park, Eunhwan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918156604276736
author Lim, Junghwan
Lee, Sungmin
Kim, Dongseok
Cheung, Wai Ting
Kim, Beomgyu
Kim, Taehwan
Lee, Haesol
Lee, Junhyeok
Oh, Dongpin
Park, Eunhwan
author_facet Lim, Junghwan
Lee, Sungmin
Kim, Dongseok
Cheung, Wai Ting
Kim, Beomgyu
Kim, Taehwan
Lee, Haesol
Lee, Junhyeok
Oh, Dongpin
Park, Eunhwan
contents The self-attention mechanism, while foundational to modern Transformer architectures, suffers from a critical inefficiency: it frequently allocates substantial attention to redundant or noisy context. Differential Attention addressed this by using subtractive attention maps for signal and noise, but its required balanced head allocation imposes rigid constraints on representational flexibility and scalability. To overcome this, we propose Grouped Differential Attention (GDA), a novel approach that introduces unbalanced head allocation between signal-preserving and noise-control groups. GDA significantly enhances signal focus by strategically assigning more heads to signal extraction and fewer to noise-control, stabilizing the latter through controlled repetition (akin to GQA). This design achieves stronger signal fidelity with minimal computational overhead. We further extend this principle to group-differentiated growth, a scalable strategy that selectively replicates only the signal-focused heads, thereby ensuring efficient capacity expansion. Through large-scale pretraining and continual training experiments, we demonstrate that moderate imbalance ratios in GDA yield substantial improvements in generalization and stability compared to symmetric baselines. Our results collectively establish that ratio-aware head allocation and selective expansion offer an effective and practical path toward designing scalable, computation-efficient Transformer architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Grouped Differential Attention
Lim, Junghwan
Lee, Sungmin
Kim, Dongseok
Cheung, Wai Ting
Kim, Beomgyu
Kim, Taehwan
Lee, Haesol
Lee, Junhyeok
Oh, Dongpin
Park, Eunhwan
Machine Learning
Artificial Intelligence
The self-attention mechanism, while foundational to modern Transformer architectures, suffers from a critical inefficiency: it frequently allocates substantial attention to redundant or noisy context. Differential Attention addressed this by using subtractive attention maps for signal and noise, but its required balanced head allocation imposes rigid constraints on representational flexibility and scalability. To overcome this, we propose Grouped Differential Attention (GDA), a novel approach that introduces unbalanced head allocation between signal-preserving and noise-control groups. GDA significantly enhances signal focus by strategically assigning more heads to signal extraction and fewer to noise-control, stabilizing the latter through controlled repetition (akin to GQA). This design achieves stronger signal fidelity with minimal computational overhead. We further extend this principle to group-differentiated growth, a scalable strategy that selectively replicates only the signal-focused heads, thereby ensuring efficient capacity expansion. Through large-scale pretraining and continual training experiments, we demonstrate that moderate imbalance ratios in GDA yield substantial improvements in generalization and stability compared to symmetric baselines. Our results collectively establish that ratio-aware head allocation and selective expansion offer an effective and practical path toward designing scalable, computation-efficient Transformer architectures.
title Grouped Differential Attention
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.06949