Sinkhorn doubly stochastic attention rank decay analysis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lapenna, Michela, Fioresi, Rita, Gharesifard, Bahman
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915926760226816
author Lapenna, Michela
Fioresi, Rita
Gharesifard, Bahman
author_facet Lapenna, Michela
Fioresi, Rita
Gharesifard, Bahman
contents The self-attention mechanism is central to the success of Transformer architectures. However, standard row-stochastic attention has been shown to suffer from significant signal degradation across layers. In particular, it can induce rank collapse, resulting in increasingly uniform token representations, as well as entropy collapse, characterized by highly concentrated attention distributions. Recent work has highlighted the benefits of doubly stochastic attention as a form of entropy regularization, promoting a more balanced attention distribution and leading to improved empirical performance. In this paper, we study rank collapse across network depth and show that doubly stochastic attention matrices normalized with Sinkhorn algorithm preserve rank more effectively than standard Softmax row-stochastic ones. As previously shown for Softmax, skip connections are crucial to mitigate rank collapse. We empirically validate this phenomenon on both sentiment analysis and image classification tasks. Moreover, we derive a theoretical bound for the pure self-attention rank decay when using Sinkhorn normalization and find that rank decays to one doubly exponentially with depth, a phenomenon that has already been shown for Softmax.
format Preprint
id arxiv_https___arxiv_org_abs_2604_07925
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Sinkhorn doubly stochastic attention rank decay analysis
Lapenna, Michela
Fioresi, Rita
Gharesifard, Bahman
Machine Learning
Artificial Intelligence
Optimization and Control
The self-attention mechanism is central to the success of Transformer architectures. However, standard row-stochastic attention has been shown to suffer from significant signal degradation across layers. In particular, it can induce rank collapse, resulting in increasingly uniform token representations, as well as entropy collapse, characterized by highly concentrated attention distributions. Recent work has highlighted the benefits of doubly stochastic attention as a form of entropy regularization, promoting a more balanced attention distribution and leading to improved empirical performance. In this paper, we study rank collapse across network depth and show that doubly stochastic attention matrices normalized with Sinkhorn algorithm preserve rank more effectively than standard Softmax row-stochastic ones. As previously shown for Softmax, skip connections are crucial to mitigate rank collapse. We empirically validate this phenomenon on both sentiment analysis and image classification tasks. Moreover, we derive a theoretical bound for the pure self-attention rank decay when using Sinkhorn normalization and find that rank decays to one doubly exponentially with depth, a phenomenon that has already been shown for Softmax.
title Sinkhorn doubly stochastic attention rank decay analysis
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2604.07925