Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix
Fuente:
arXiv
Saved in:
| Main Authors: | Hayase, Tomohiro, Collins, Benoît, Karakida, Ryo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention
by: Hayase, Tomohiro, et al.
Published: (2026)
by: Hayase, Tomohiro, et al.
Published: (2026)
Free Random Projection for In-Context Reinforcement Learning
by: Hayase, Tomohiro, et al.
Published: (2025)
by: Hayase, Tomohiro, et al.
Published: (2025)
Understanding MLP-Mixer as a Wide and Sparse MLP
by: Hayase, Tomohiro, et al.
Published: (2023)
by: Hayase, Tomohiro, et al.
Published: (2023)
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
by: Tomihari, Akiyoshi, et al.
Published: (2025)
by: Tomihari, Akiyoshi, et al.
Published: (2025)
Asymptotic Gaussian Fluctuations of Eigenvectors in Spectral Clustering
by: Lebeau, Hugo, et al.
Published: (2024)
by: Lebeau, Hugo, et al.
Published: (2024)
Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor Programs
by: Sakai, Mana, et al.
Published: (2025)
by: Sakai, Mana, et al.
Published: (2025)
Non-Asymptotic Analysis of Data Augmentation for Precision Matrix Estimation
by: Morisset, Lucas, et al.
Published: (2025)
by: Morisset, Lucas, et al.
Published: (2025)
Asymptotic spectrum of weighted sample covariance: another proof of spectrum convergence
by: Oriol, Benoit
Published: (2024)
by: Oriol, Benoit
Published: (2024)
Asymptotic non-linear shrinkage and eigenvector overlap for weighted sample covariance
by: Oriol, Benoit
Published: (2024)
by: Oriol, Benoit
Published: (2024)
High-dimensional Asymptotics of Langevin Dynamics in Spiked Matrix Models
by: Liang, Tengyuan, et al.
Published: (2022)
by: Liang, Tengyuan, et al.
Published: (2022)
Self-attention Networks Localize When QK-eigenspectrum Concentrates
by: Bao, Han, et al.
Published: (2024)
by: Bao, Han, et al.
Published: (2024)
Gaussian Processes and Reproducing Kernels: Connections and Equivalences
by: Kanagawa, Motonobu, et al.
Published: (2025)
by: Kanagawa, Motonobu, et al.
Published: (2025)
On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width
by: Ishikawa, Satoki, et al.
Published: (2023)
by: Ishikawa, Satoki, et al.
Published: (2023)
Analysis of a multi-target linear shrinkage covariance estimator
by: Oriol, Benoit
Published: (2024)
by: Oriol, Benoit
Published: (2024)
Convergent Stochastic Training of Attention and Understanding LoRA
by: Sun, Zhengkai, et al.
Published: (2026)
by: Sun, Zhengkai, et al.
Published: (2026)
Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
by: Zucker, Shai, et al.
Published: (2025)
by: Zucker, Shai, et al.
Published: (2025)
Stochastic Clock Attention for Aligning Continuous and Ordered Sequences
by: Soh, Hyungjoon, et al.
Published: (2025)
by: Soh, Hyungjoon, et al.
Published: (2025)
Optimal Layer Selection for Latent Data Augmentation
by: Takase, Tomoumi, et al.
Published: (2024)
by: Takase, Tomoumi, et al.
Published: (2024)
Spectral Estimators for Multi-Index Models: Precise Asymptotics and Optimal Weak Recovery
by: Kovačević, Filip, et al.
Published: (2025)
by: Kovačević, Filip, et al.
Published: (2025)
Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework
by: Chen, Zaiwei, et al.
Published: (2026)
by: Chen, Zaiwei, et al.
Published: (2026)
Fundamental Limits of Matrix Sensing: Exact Asymptotics, Universality, and Applications
by: Xu, Yizhou, et al.
Published: (2025)
by: Xu, Yizhou, et al.
Published: (2025)
Matrix Denoising with Doubly Heteroscedastic Noise: Fundamental Limits and Optimal Spectral Methods
by: Zhang, Yihan, et al.
Published: (2024)
by: Zhang, Yihan, et al.
Published: (2024)
Continual Learning in Modern Hopfield Networks with an Application to Diffusion Models
by: Takeda, Ken, et al.
Published: (2026)
by: Takeda, Ken, et al.
Published: (2026)
Permutation-Invariant Spectral Learning via Dyson Diffusion
by: Schwarz, Tassilo, et al.
Published: (2025)
by: Schwarz, Tassilo, et al.
Published: (2025)
A Random Matrix Approach to Low-Multilinear-Rank Tensor Approximation
by: Lebeau, Hugo, et al.
Published: (2024)
by: Lebeau, Hugo, et al.
Published: (2024)
Gaussian Approximation for Asynchronous Q-learning
by: Rubtsov, Artemy, et al.
Published: (2026)
by: Rubtsov, Artemy, et al.
Published: (2026)
Obstacle-aware Gaussian Process Regression
by: Shrivastava, Gaurav
Published: (2024)
by: Shrivastava, Gaurav
Published: (2024)
Performance Gaps in Multi-view Clustering under the Nested Matrix-Tensor Model
by: Lebeau, Hugo, et al.
Published: (2024)
by: Lebeau, Hugo, et al.
Published: (2024)
On the Equivalence of Random Network Distillation, Deep Ensembles, and Bayesian Inference
by: Zanger, Moritz A., et al.
Published: (2026)
by: Zanger, Moritz A., et al.
Published: (2026)
Spectral alignment of stochastic gradient descent for high-dimensional classification tasks
by: Arous, Gerard Ben, et al.
Published: (2023)
by: Arous, Gerard Ben, et al.
Published: (2023)
Mixing of the No-U-Turn Sampler and the Geometry of Gaussian Concentration
by: Bou-Rabee, Nawaf, et al.
Published: (2024)
by: Bou-Rabee, Nawaf, et al.
Published: (2024)
Resolving Node Identifiability in Graph Neural Processes via Laplacian Spectral Encodings
by: Yan, Zimo, et al.
Published: (2025)
by: Yan, Zimo, et al.
Published: (2025)
Random ReLU Neural Networks as Non-Gaussian Processes
by: Parhi, Rahul, et al.
Published: (2024)
by: Parhi, Rahul, et al.
Published: (2024)
Large-width functional asymptotics for deep Gaussian neural networks
by: Bracale, Daniele, et al.
Published: (2021)
by: Bracale, Daniele, et al.
Published: (2021)
A Gaussian Comparison Theorem for Training Dynamics in Machine Learning
by: Panahi, Ashkan
Published: (2026)
by: Panahi, Ashkan
Published: (2026)
WeSpeR: Computing non-linear shrinkage formulas for the weighted sample covariance
by: Oriol, Benoit
Published: (2024)
by: Oriol, Benoit
Published: (2024)
Eigen-Spike Emergence and Quadratic Equivalents for Conjugate Kernels on Nonlinearly Separable Data
by: Cranston, Collin, et al.
Published: (2026)
by: Cranston, Collin, et al.
Published: (2026)
Asymptotic convexity of wide and shallow neural networks
by: Borkar, Vivek, et al.
Published: (2025)
by: Borkar, Vivek, et al.
Published: (2025)
Rethinking Nonlinearity: Trainable Gaussian Mixture Modules for Modern Neural Architectures
by: Lu, Weiguo, et al.
Published: (2025)
by: Lu, Weiguo, et al.
Published: (2025)
Differentiable Expectation-Maximisation and Applications to Gaussian Mixture Model Optimal Transport
by: Boïté, Samuel, et al.
Published: (2025)
by: Boïté, Samuel, et al.
Published: (2025)
Similar Items
-
A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention
by: Hayase, Tomohiro, et al.
Published: (2026) -
Free Random Projection for In-Context Reinforcement Learning
by: Hayase, Tomohiro, et al.
Published: (2025) -
Understanding MLP-Mixer as a Wide and Sparse MLP
by: Hayase, Tomohiro, et al.
Published: (2023) -
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
by: Tomihari, Akiyoshi, et al.
Published: (2025) -
Asymptotic Gaussian Fluctuations of Eigenvectors in Spectral Clustering
by: Lebeau, Hugo, et al.
Published: (2024)