Statistical Advantage of Softmax Attention: Insights from Single-Location Regression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duranthon, O., Marion, P., Boyer, C., Loureiro, B., Zdeborová, L.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912927294947328
author Duranthon, O.
Marion, P.
Boyer, C.
Loureiro, B.
Zdeborová, L.
author_facet Duranthon, O.
Marion, P.
Boyer, C.
Loureiro, B.
Zdeborová, L.
contents Large language models rely on attention mechanisms with a softmax activation. Yet the dominance of softmax over alternatives (e.g., component-wise or linear) remains poorly understood, and many theoretical works have focused on the easier-to-analyze linearized attention. In this work, we address this gap through a principled study of the single-location regression task, where the output depends on a linear transformation of a single input token at a random location. Building on ideas from statistical physics, we develop an analysis of attention-based predictors in the high-dimensional limit, where generalization performance is captured by a small set of order parameters. At the population level, we show that softmax achieves the Bayes risk, whereas linear attention fundamentally falls short. We then examine other activation functions to identify which properties are necessary for optimal performance. Finally, we analyze the finite-sample regime: we provide an asymptotic characterization of the test error and show that, while softmax is no longer Bayes-optimal, it consistently outperforms linear attention. We discuss the connection with optimization by gradient-based algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21936
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
Duranthon, O.
Marion, P.
Boyer, C.
Loureiro, B.
Zdeborová, L.
Machine Learning
Disordered Systems and Neural Networks
Large language models rely on attention mechanisms with a softmax activation. Yet the dominance of softmax over alternatives (e.g., component-wise or linear) remains poorly understood, and many theoretical works have focused on the easier-to-analyze linearized attention. In this work, we address this gap through a principled study of the single-location regression task, where the output depends on a linear transformation of a single input token at a random location. Building on ideas from statistical physics, we develop an analysis of attention-based predictors in the high-dimensional limit, where generalization performance is captured by a small set of order parameters. At the population level, we show that softmax achieves the Bayes risk, whereas linear attention fundamentally falls short. We then examine other activation functions to identify which properties are necessary for optimal performance. Finally, we analyze the finite-sample regime: we provide an asymptotic characterization of the test error and show that, while softmax is no longer Bayes-optimal, it consistently outperforms linear attention. We discuss the connection with optimization by gradient-based algorithms.
title Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
topic Machine Learning
Disordered Systems and Neural Networks
url https://arxiv.org/abs/2509.21936