Training Dynamics of Softmax Self-Attention: Fast Global Convergence via Preconditioning
Fuente:
arXiv
Salvato in:
| Autori principali: | Goel, Gautam, Soltanolkotabi, Mahdi, Bartlett, Peter |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can a Transformer Represent a Kalman Filter?
di: Goel, Gautam, et al.
Pubblicazione: (2023)
di: Goel, Gautam, et al.
Pubblicazione: (2023)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
di: Chen, Siyu, et al.
Pubblicazione: (2024)
di: Chen, Siyu, et al.
Pubblicazione: (2024)
Fast Convergence of Softmax Policy Mirror Ascent
di: Asad, Reza, et al.
Pubblicazione: (2024)
di: Asad, Reza, et al.
Pubblicazione: (2024)
Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable Rewards
di: Heckel, Reinhard, et al.
Pubblicazione: (2026)
di: Heckel, Reinhard, et al.
Pubblicazione: (2026)
Theoretical Insights into Overparameterized Models in Multi-Task and Replay-Based Continual Learning
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2024)
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2024)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
di: Lin, Max Qiushi, et al.
Pubblicazione: (2025)
di: Lin, Max Qiushi, et al.
Pubblicazione: (2025)
In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention
di: He, Jianliang, et al.
Pubblicazione: (2025)
di: He, Jianliang, et al.
Pubblicazione: (2025)
On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
di: Labbi, Safwan, et al.
Pubblicazione: (2025)
di: Labbi, Safwan, et al.
Pubblicazione: (2025)
Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models
di: Fabian, Zalan, et al.
Pubblicazione: (2023)
di: Fabian, Zalan, et al.
Pubblicazione: (2023)
Convergence Rates for Softmax Gating Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
NDCG-Consistent Softmax Approximation with Accelerated Convergence
di: Pu, Yuanhao, et al.
Pubblicazione: (2025)
di: Pu, Yuanhao, et al.
Pubblicazione: (2025)
Softmax Linear Attention: Reclaiming Global Competition
di: Xu, Mingwei, et al.
Pubblicazione: (2026)
di: Xu, Mingwei, et al.
Pubblicazione: (2026)
Test-Time Training Provably Improves Transformers as In-context Learners
di: Gozeten, Halil Alperen, et al.
Pubblicazione: (2025)
di: Gozeten, Halil Alperen, et al.
Pubblicazione: (2025)
FoNE: Precise Single-Token Number Embeddings via Fourier Features
di: Zhou, Tianyi, et al.
Pubblicazione: (2025)
di: Zhou, Tianyi, et al.
Pubblicazione: (2025)
Universal Approximation with Softmax Attention
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2025)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2025)
Learning to Recall with Transformers Beyond Orthogonal Embeddings
di: Vural, Nuri Mert, et al.
Pubblicazione: (2026)
di: Vural, Nuri Mert, et al.
Pubblicazione: (2026)
Why Softmax Attention Outperforms Linear Attention
di: Deng, Yichuan, et al.
Pubblicazione: (2023)
di: Deng, Yichuan, et al.
Pubblicazione: (2023)
MosaicMRI: A Diverse Dataset and Benchmark for Raw Musculoskeletal MRI
di: Arguello, Paula, et al.
Pubblicazione: (2026)
di: Arguello, Paula, et al.
Pubblicazione: (2026)
LUCID: Attention with Preconditioned Representations
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
Provable Multi-Task Representation Learning by Two-Layer ReLU Neural Networks
di: Collins, Liam, et al.
Pubblicazione: (2023)
di: Collins, Liam, et al.
Pubblicazione: (2023)
Scalable-Softmax Is Superior for Attention
di: Nakanishi, Ken M.
Pubblicazione: (2025)
di: Nakanishi, Ken M.
Pubblicazione: (2025)
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
di: Zhang, Ruiqi, et al.
Pubblicazione: (2025)
di: Zhang, Ruiqi, et al.
Pubblicazione: (2025)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
di: Yan, Fanqi, et al.
Pubblicazione: (2025)
di: Yan, Fanqi, et al.
Pubblicazione: (2025)
On the Invariants of Softmax Attention
di: Lee, Wonsuk
Pubblicazione: (2026)
di: Lee, Wonsuk
Pubblicazione: (2026)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
di: Zuhri, Zayd M. K., et al.
Pubblicazione: (2025)
di: Zuhri, Zayd M. K., et al.
Pubblicazione: (2025)
Preconditioned Attention: Enhancing Efficiency in Transformers
di: Saratchandran, Hemanth
Pubblicazione: (2026)
di: Saratchandran, Hemanth
Pubblicazione: (2026)
Serpent: Scalable and Efficient Image Restoration via Multi-scale Structured State Space Models
di: Sepehri, Mohammad Shahab, et al.
Pubblicazione: (2024)
di: Sepehri, Mohammad Shahab, et al.
Pubblicazione: (2024)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
di: Nishikawa, Naoki, et al.
Pubblicazione: (2025)
di: Nishikawa, Naoki, et al.
Pubblicazione: (2025)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
di: Kovačević, Filip, et al.
Pubblicazione: (2026)
di: Kovačević, Filip, et al.
Pubblicazione: (2026)
Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity
di: Shi, Zhongjie, et al.
Pubblicazione: (2026)
di: Shi, Zhongjie, et al.
Pubblicazione: (2026)
Softmax Attention with Constant Cost per Token
di: Heinsen, Franz A.
Pubblicazione: (2024)
di: Heinsen, Franz A.
Pubblicazione: (2024)
LMC: Fast Training of GNNs via Subgraph Sampling with Provable Convergence
di: Shi, Zhihao, et al.
Pubblicazione: (2023)
di: Shi, Zhihao, et al.
Pubblicazione: (2023)
ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization
di: Gan, Haosheng, et al.
Pubblicazione: (2025)
di: Gan, Haosheng, et al.
Pubblicazione: (2025)
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
di: Mongaras, Gabriel, et al.
Pubblicazione: (2025)
di: Mongaras, Gabriel, et al.
Pubblicazione: (2025)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
di: Kuang, Yilun, et al.
Pubblicazione: (2025)
di: Kuang, Yilun, et al.
Pubblicazione: (2025)
CrispEdit: Low-Curvature Projections for Scalable Non-Destructive LLM Editing
di: Ikram, Zarif, et al.
Pubblicazione: (2026)
di: Ikram, Zarif, et al.
Pubblicazione: (2026)
Emergence and Evolution of Interpretable Concepts in Diffusion Models
di: Tinaz, Berk, et al.
Pubblicazione: (2025)
di: Tinaz, Berk, et al.
Pubblicazione: (2025)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
di: Rezazadeh, Navid, et al.
Pubblicazione: (2026)
di: Rezazadeh, Navid, et al.
Pubblicazione: (2026)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
di: Zhang, Michael, et al.
Pubblicazione: (2024)
di: Zhang, Michael, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Can a Transformer Represent a Kalman Filter?
di: Goel, Gautam, et al.
Pubblicazione: (2023) -
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
di: Chen, Siyu, et al.
Pubblicazione: (2024) -
Fast Convergence of Softmax Policy Mirror Ascent
di: Asad, Reza, et al.
Pubblicazione: (2024) -
Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable Rewards
di: Heckel, Reinhard, et al.
Pubblicazione: (2026) -
Theoretical Insights into Overparameterized Models in Multi-Task and Replay-Based Continual Learning
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2024)