Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity
Fuente:
arXiv
Guardado en:
| Autores principales: | Shi, Zhongjie, Liao, Wenjing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data
por: Havrilla, Alex, et al.
Publicado: (2024)
por: Havrilla, Alex, et al.
Publicado: (2024)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
por: Lin, Max Qiushi, et al.
Publicado: (2025)
por: Lin, Max Qiushi, et al.
Publicado: (2025)
Theory of Decentralized Robust Kernel-Based Learning
por: Yu, Zhan, et al.
Publicado: (2025)
por: Yu, Zhan, et al.
Publicado: (2025)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
por: Rezazadeh, Navid, et al.
Publicado: (2026)
por: Rezazadeh, Navid, et al.
Publicado: (2026)
Universal Approximation with Softmax Attention
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2025)
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2025)
NDCG-Consistent Softmax Approximation with Accelerated Convergence
por: Pu, Yuanhao, et al.
Publicado: (2025)
por: Pu, Yuanhao, et al.
Publicado: (2025)
Transformers for Learning on Noisy and Task-Level Manifolds: Approximation and Generalization Insights
por: Shen, Zhaiming, et al.
Publicado: (2025)
por: Shen, Zhaiming, et al.
Publicado: (2025)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
por: Gonsior, Julius, et al.
Publicado: (2022)
por: Gonsior, Julius, et al.
Publicado: (2022)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
por: Xie, Zixuan, et al.
Publicado: (2026)
por: Xie, Zixuan, et al.
Publicado: (2026)
Online Continual Learning via Logit Adjusted Softmax
por: Huang, Zhehao, et al.
Publicado: (2023)
por: Huang, Zhehao, et al.
Publicado: (2023)
Training Dynamics of Softmax Self-Attention: Fast Global Convergence via Preconditioning
por: Goel, Gautam, et al.
Publicado: (2026)
por: Goel, Gautam, et al.
Publicado: (2026)
Deep Neural Networks are Adaptive to Function Regularity and Data Distribution in Approximation and Estimation
por: Liu, Hao, et al.
Publicado: (2024)
por: Liu, Hao, et al.
Publicado: (2024)
Partition of Unity Neural Networks for Interpretable Classification with Explicit Class Regions
por: Aldroubi, Akram
Publicado: (2026)
por: Aldroubi, Akram
Publicado: (2026)
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
por: Lv, Qi, et al.
Publicado: (2025)
por: Lv, Qi, et al.
Publicado: (2025)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
por: Collins, Liam, et al.
Publicado: (2024)
por: Collins, Liam, et al.
Publicado: (2024)
Softmax-free Linear Transformers
por: Lu, Jiachen, et al.
Publicado: (2022)
por: Lu, Jiachen, et al.
Publicado: (2022)
Approximation Theory for Lipschitz Continuous Transformers
por: Furuya, Takashi, et al.
Publicado: (2026)
por: Furuya, Takashi, et al.
Publicado: (2026)
Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer
por: Hsu, Alexander, et al.
Publicado: (2026)
por: Hsu, Alexander, et al.
Publicado: (2026)
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
por: Chou, Yuhong, et al.
Publicado: (2024)
por: Chou, Yuhong, et al.
Publicado: (2024)
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2023)
por: Nguyen, Huy, et al.
Publicado: (2023)
Adaptive Sampled Softmax with Inverted Multi-Index: Methods, Theory and Applications
por: Chen, Jin, et al.
Publicado: (2025)
por: Chen, Jin, et al.
Publicado: (2025)
Softmax Transformers are Turing-Complete
por: Jiang, Hongjian, et al.
Publicado: (2025)
por: Jiang, Hongjian, et al.
Publicado: (2025)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
por: Peltekis, Christodoulos, et al.
Publicado: (2024)
por: Peltekis, Christodoulos, et al.
Publicado: (2024)
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
por: Saratchandran, Hemanth, et al.
Publicado: (2024)
por: Saratchandran, Hemanth, et al.
Publicado: (2024)
On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
por: Labbi, Safwan, et al.
Publicado: (2025)
por: Labbi, Safwan, et al.
Publicado: (2025)
Transformers Meet In-Context Learning: A Universal Approximation Theory
por: Li, Gen, et al.
Publicado: (2025)
por: Li, Gen, et al.
Publicado: (2025)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
por: İslamoğlu, Gamze, et al.
Publicado: (2023)
por: İslamoğlu, Gamze, et al.
Publicado: (2023)
BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling
por: Ye, Zisheng, et al.
Publicado: (2026)
por: Ye, Zisheng, et al.
Publicado: (2026)
Is Softmax Loss All You Need? A Principled Analysis of Softmax-family Loss
por: Pu, Yuanhao, et al.
Publicado: (2026)
por: Pu, Yuanhao, et al.
Publicado: (2026)
Distributed Gradient Descent for Functional Learning
por: Yu, Zhan, et al.
Publicado: (2023)
por: Yu, Zhan, et al.
Publicado: (2023)
$ε$-Softmax: Approximating One-Hot Vectors for Mitigating Label Noise
por: Wang, Jialiang, et al.
Publicado: (2025)
por: Wang, Jialiang, et al.
Publicado: (2025)
Decorr: Environment Partitioning for Invariant Learning and OOD Generalization
por: Liao, Yufan, et al.
Publicado: (2022)
por: Liao, Yufan, et al.
Publicado: (2022)
Forgetting Transformer: Softmax Attention with a Forget Gate
por: Lin, Zhixuan, et al.
Publicado: (2025)
por: Lin, Zhixuan, et al.
Publicado: (2025)
N-Adaptive Ritz Method: A Neural Network Enriched Partition of Unity for Boundary Value Problems
por: Baek, Jonghyuk, et al.
Publicado: (2024)
por: Baek, Jonghyuk, et al.
Publicado: (2024)
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
por: Ran-Milo, Yuval
Publicado: (2026)
por: Ran-Milo, Yuval
Publicado: (2026)
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
por: Dragutinović, Sara, et al.
Publicado: (2025)
por: Dragutinović, Sara, et al.
Publicado: (2025)
Softmax Linear Attention: Reclaiming Global Competition
por: Xu, Mingwei, et al.
Publicado: (2026)
por: Xu, Mingwei, et al.
Publicado: (2026)
Partition of Unity Physics-Informed Neural Networks (POU-PINNs): An Unsupervised Framework for Physics-Informed Domain Decomposition and Mixtures of Experts
por: Rodriguez, Arturo, et al.
Publicado: (2024)
por: Rodriguez, Arturo, et al.
Publicado: (2024)
GLOP: Learning Global Partition and Local Construction for Solving Large-scale Routing Problems in Real-time
por: Ye, Haoran, et al.
Publicado: (2023)
por: Ye, Haoran, et al.
Publicado: (2023)
Local to Global: Learning Dynamics and Effect of Initialization for Transformers
por: Makkuva, Ashok Vardhan, et al.
Publicado: (2024)
por: Makkuva, Ashok Vardhan, et al.
Publicado: (2024)
Ejemplares similares
-
Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data
por: Havrilla, Alex, et al.
Publicado: (2024) -
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
por: Lin, Max Qiushi, et al.
Publicado: (2025) -
Theory of Decentralized Robust Kernel-Based Learning
por: Yu, Zhan, et al.
Publicado: (2025) -
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
por: Rezazadeh, Navid, et al.
Publicado: (2026) -
Universal Approximation with Softmax Attention
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2025)