Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915815085834240 |
|---|---|
| author | Zucker, Shai Wang, Xiong Lu, Fei Seroussi, Inbar |
| author_facet | Zucker, Shai Wang, Xiong Lu, Fei Seroussi, Inbar |
| contents | We study the convergence rate of learning pairwise interactions in single-layer attention-style models, where tokens interact through a weight matrix and a nonlinear activation function. We prove that the minimax rate is $M^{-\frac{2β}{2β+1}}$, where $M$ is the sample size and $β$ is the Hölder smoothness of the activation function. Importantly, this rate is independent of the embedding dimension $d$, the number of tokens $N$, and the rank $r$ of the weight matrix, provided that $rd \le (M/\log M)^{\frac{1}{2β+1}}$. These results highlight a fundamental statistical efficiency of attention-style models, even when the weight matrix and activation are not separately identifiable, and provide a theoretical understanding of attention mechanisms and guidance on training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_11789 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Minimax Rates for Learning Pairwise Interactions in Attention-Style Models Zucker, Shai Wang, Xiong Lu, Fei Seroussi, Inbar Machine Learning Probability Statistics Theory We study the convergence rate of learning pairwise interactions in single-layer attention-style models, where tokens interact through a weight matrix and a nonlinear activation function. We prove that the minimax rate is $M^{-\frac{2β}{2β+1}}$, where $M$ is the sample size and $β$ is the Hölder smoothness of the activation function. Importantly, this rate is independent of the embedding dimension $d$, the number of tokens $N$, and the rank $r$ of the weight matrix, provided that $rd \le (M/\log M)^{\frac{1}{2β+1}}$. These results highlight a fundamental statistical efficiency of attention-style models, even when the weight matrix and activation are not separately identifiable, and provide a theoretical understanding of attention mechanisms and guidance on training. |
| title | Minimax Rates for Learning Pairwise Interactions in Attention-Style Models |
| topic | Machine Learning Probability Statistics Theory |
| url | https://arxiv.org/abs/2510.11789 |