Minimax Rates for Learning Pairwise Interactions in Attention-Style Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zucker, Shai, Wang, Xiong, Lu, Fei, Seroussi, Inbar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915815085834240
author Zucker, Shai
Wang, Xiong
Lu, Fei
Seroussi, Inbar
author_facet Zucker, Shai
Wang, Xiong
Lu, Fei
Seroussi, Inbar
contents We study the convergence rate of learning pairwise interactions in single-layer attention-style models, where tokens interact through a weight matrix and a nonlinear activation function. We prove that the minimax rate is $M^{-\frac{2β}{2β+1}}$, where $M$ is the sample size and $β$ is the Hölder smoothness of the activation function. Importantly, this rate is independent of the embedding dimension $d$, the number of tokens $N$, and the rank $r$ of the weight matrix, provided that $rd \le (M/\log M)^{\frac{1}{2β+1}}$. These results highlight a fundamental statistical efficiency of attention-style models, even when the weight matrix and activation are not separately identifiable, and provide a theoretical understanding of attention mechanisms and guidance on training.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11789
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
Zucker, Shai
Wang, Xiong
Lu, Fei
Seroussi, Inbar
Machine Learning
Probability
Statistics Theory
We study the convergence rate of learning pairwise interactions in single-layer attention-style models, where tokens interact through a weight matrix and a nonlinear activation function. We prove that the minimax rate is $M^{-\frac{2β}{2β+1}}$, where $M$ is the sample size and $β$ is the Hölder smoothness of the activation function. Importantly, this rate is independent of the embedding dimension $d$, the number of tokens $N$, and the rank $r$ of the weight matrix, provided that $rd \le (M/\log M)^{\frac{1}{2β+1}}$. These results highlight a fundamental statistical efficiency of attention-style models, even when the weight matrix and activation are not separately identifiable, and provide a theoretical understanding of attention mechanisms and guidance on training.
title Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
topic Machine Learning
Probability
Statistics Theory
url https://arxiv.org/abs/2510.11789