Limitations of Normalization in Attention Mechanism

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mudarisov, Timur, Burtsev, Mikhail, Petrova, Tatiana, State, Radu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915564057788416
author Mudarisov, Timur
Burtsev, Mikhail
Petrova, Tatiana
State, Radu
author_facet Mudarisov, Timur
Burtsev, Mikhail
Petrova, Tatiana
State, Radu
contents This paper investigates the limitations of the normalization in attention mechanisms. We begin with a theoretical framework that enables the identification of the model's selective ability and the geometric separation involved in token selection. Our analysis includes explicit bounds on distances and separation criteria for token vectors under softmax scaling. Through experiments with pre-trained GPT-2 model, we empirically validate our theoretical results and analyze key behaviors of the attention mechanism. Notably, we demonstrate that as the number of selected tokens increases, the model's ability to distinguish informative tokens declines, often converging toward a uniform selection pattern. We also show that gradient sensitivity under softmax normalization presents challenges during training, especially at low temperature settings. These findings advance current understanding of softmax-based attention mechanism and motivate the need for more robust normalization and selection strategies in future attention architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Limitations of Normalization in Attention Mechanism
Mudarisov, Timur
Burtsev, Mikhail
Petrova, Tatiana
State, Radu
Machine Learning
Artificial Intelligence
Computation and Language
This paper investigates the limitations of the normalization in attention mechanisms. We begin with a theoretical framework that enables the identification of the model's selective ability and the geometric separation involved in token selection. Our analysis includes explicit bounds on distances and separation criteria for token vectors under softmax scaling. Through experiments with pre-trained GPT-2 model, we empirically validate our theoretical results and analyze key behaviors of the attention mechanism. Notably, we demonstrate that as the number of selected tokens increases, the model's ability to distinguish informative tokens declines, often converging toward a uniform selection pattern. We also show that gradient sensitivity under softmax normalization presents challenges during training, especially at low temperature settings. These findings advance current understanding of softmax-based attention mechanism and motivate the need for more robust normalization and selection strategies in future attention architectures.
title Limitations of Normalization in Attention Mechanism
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.17821