Self-attention Networks Localize When QK-eigenspectrum Concentrates
Fuente:
arXiv
Saved in:
| Main Authors: | Bao, Han, Hataya, Ryuichiro, Karakida, Ryo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Noncommutative $C^*$-algebra Net: Learning Neural Networks with Powerful Product Structure in $C^*$-algebra
by: Hataya, Ryuichiro, et al.
Published: (2023)
by: Hataya, Ryuichiro, et al.
Published: (2023)
Glocal Hypergradient Estimation with Koopman Operator
by: Hataya, Ryuichiro, et al.
Published: (2024)
by: Hataya, Ryuichiro, et al.
Published: (2024)
Quantum Circuit $C^*$-algebra Net
by: Hashimoto, Yuka, et al.
Published: (2024)
by: Hashimoto, Yuka, et al.
Published: (2024)
Automatic Domain Adaptation by Transformers in In-Context Learning
by: Hataya, Ryuichiro, et al.
Published: (2024)
by: Hataya, Ryuichiro, et al.
Published: (2024)
Provable Target Sample Complexity Improvements as Pre-Trained Models Scale
by: Fukuchi, Kazuto, et al.
Published: (2026)
by: Fukuchi, Kazuto, et al.
Published: (2026)
Provable Data Scaling Law for Meta Learning via Complexity Minimization
by: Fukuchi, Kazuto, et al.
Published: (2026)
by: Fukuchi, Kazuto, et al.
Published: (2026)
Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target Propagation
by: Ishikawa, Satoki, et al.
Published: (2024)
by: Ishikawa, Satoki, et al.
Published: (2024)
Understanding MLP-Mixer as a Wide and Sparse MLP
by: Hayase, Tomohiro, et al.
Published: (2023)
by: Hayase, Tomohiro, et al.
Published: (2023)
On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width
by: Ishikawa, Satoki, et al.
Published: (2023)
by: Ishikawa, Satoki, et al.
Published: (2023)
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
by: Tomihari, Akiyoshi, et al.
Published: (2025)
by: Tomihari, Akiyoshi, et al.
Published: (2025)
Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix
by: Hayase, Tomohiro, et al.
Published: (2025)
by: Hayase, Tomohiro, et al.
Published: (2025)
Continual Learning in Modern Hopfield Networks with an Application to Diffusion Models
by: Takeda, Ken, et al.
Published: (2026)
by: Takeda, Ken, et al.
Published: (2026)
The Impact of Anisotropic Covariance Structure on the Training Dynamics and Generalization Error of Linear Networks
by: Watanabe, Taishi, et al.
Published: (2026)
by: Watanabe, Taishi, et al.
Published: (2026)
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
by: Naganuma, Hiroki, et al.
Published: (2023)
by: Naganuma, Hiroki, et al.
Published: (2023)
Optimal Layer Selection for Latent Data Augmentation
by: Takase, Tomoumi, et al.
Published: (2024)
by: Takase, Tomoumi, et al.
Published: (2024)
A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention
by: Hayase, Tomohiro, et al.
Published: (2026)
by: Hayase, Tomohiro, et al.
Published: (2026)
Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor Programs
by: Sakai, Mana, et al.
Published: (2025)
by: Sakai, Mana, et al.
Published: (2025)
Hierarchical Associative Memory, Parallelized MLP-Mixer, and Symmetry Breaking
by: Karakida, Ryo, et al.
Published: (2024)
by: Karakida, Ryo, et al.
Published: (2024)
Evaluating Effects of Augmented SELFIES for Molecular Understanding Using QK-LSTM
by: Beaudoin, Collin, et al.
Published: (2025)
by: Beaudoin, Collin, et al.
Published: (2025)
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
by: Grzywaczewski, Jakub, et al.
Published: (2026)
by: Grzywaczewski, Jakub, et al.
Published: (2026)
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Self-attention Dual Embedding for Graphs with Heterophily
by: Lai, Yurui, et al.
Published: (2023)
by: Lai, Yurui, et al.
Published: (2023)
Learning Granger Causality from Instance-wise Self-attentive Hawkes Processes
by: Wu, Dongxia, et al.
Published: (2024)
by: Wu, Dongxia, et al.
Published: (2024)
When Test-Time Adaptation Meets Self-Supervised Models
by: Han, Jisu, et al.
Published: (2025)
by: Han, Jisu, et al.
Published: (2025)
When Graph Neural Networks Meet Dynamic Mode Decomposition
by: Shi, Dai, et al.
Published: (2024)
by: Shi, Dai, et al.
Published: (2024)
Self-attention-based Diffusion Model for Time-series Imputation in Partial Blackout Scenarios
by: Islam, Mohammad Rafid Ul, et al.
Published: (2025)
by: Islam, Mohammad Rafid Ul, et al.
Published: (2025)
Easy attention: A simple attention mechanism for temporal predictions with transformers
by: Sanchis-Agudo, Marcial, et al.
Published: (2023)
by: Sanchis-Agudo, Marcial, et al.
Published: (2023)
What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Reorganizing attention-space geometry with expressive attention
by: Gros, Claudius
Published: (2024)
by: Gros, Claudius
Published: (2024)
Pruning Self-attentions into Convolutional Layers in Single Path
by: He, Haoyu, et al.
Published: (2021)
by: He, Haoyu, et al.
Published: (2021)
Feature Normalization Prevents Collapse of Non-contrastive Learning Dynamics
by: Bao, Han
Published: (2023)
by: Bao, Han
Published: (2023)
Controlling changes to attention logits
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Self-attention as an attractor network: transient memories without backpropagation
by: D'Amico, Francesco, et al.
Published: (2024)
by: D'Amico, Francesco, et al.
Published: (2024)
Neuronal Self-Adaptation Enhances Capacity and Robustness of Representation in Spiking Neural Networks
by: Yang, Zhuobin, et al.
Published: (2026)
by: Yang, Zhuobin, et al.
Published: (2026)
Approximation of relation functions and attention mechanisms
by: Altabaa, Awni, et al.
Published: (2024)
by: Altabaa, Awni, et al.
Published: (2024)
MUSE-Net: Missingness-aware mUlti-branching Self-attention Encoder for Irregular Longitudinal Electronic Health Records
by: Wang, Zekai, et al.
Published: (2024)
by: Wang, Zekai, et al.
Published: (2024)
Poly-attention: a general scheme for higher-order self-attention
by: Chakrabarti, Sayak, et al.
Published: (2026)
by: Chakrabarti, Sayak, et al.
Published: (2026)
Self-attentive Transformer for Fast and Accurate Postprocessing of Temperature and Wind Speed Forecasts
by: Van Poecke, Aaron, et al.
Published: (2024)
by: Van Poecke, Aaron, et al.
Published: (2024)
Treatment Effect Estimation with Differentiated Networked Effect on Graph Data
by: Lin, Xiaofeng, et al.
Published: (2026)
by: Lin, Xiaofeng, et al.
Published: (2026)
When Are Bias-Free ReLU Networks Effectively Linear Networks?
by: Zhang, Yedi, et al.
Published: (2024)
by: Zhang, Yedi, et al.
Published: (2024)
Similar Items
-
Noncommutative $C^*$-algebra Net: Learning Neural Networks with Powerful Product Structure in $C^*$-algebra
by: Hataya, Ryuichiro, et al.
Published: (2023) -
Glocal Hypergradient Estimation with Koopman Operator
by: Hataya, Ryuichiro, et al.
Published: (2024) -
Quantum Circuit $C^*$-algebra Net
by: Hashimoto, Yuka, et al.
Published: (2024) -
Automatic Domain Adaptation by Transformers in In-Context Learning
by: Hataya, Ryuichiro, et al.
Published: (2024) -
Provable Target Sample Complexity Improvements as Pre-Trained Models Scale
by: Fukuchi, Kazuto, et al.
Published: (2026)