Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Jiao, Yuling, Lai, Yanming, Wang, Yang, Yan, Bokai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Randomized Neural Networks with Locally Activation Function: Theory and Algorithm for Solving PDEs
by: Bi, Ran, et al.
Published: (2026)
by: Bi, Ran, et al.
Published: (2026)
Approximation bounds for norm constrained deep neural networks
by: Maiale, Francesco Paolo, et al.
Published: (2025)
by: Maiale, Francesco Paolo, et al.
Published: (2025)
VORT: Adaptive Power-Law Memory for NLP Transformers
by: Mlaiki, Nabil
Published: (2026)
by: Mlaiki, Nabil
Published: (2026)
Covering Numbers for Deep ReLU Networks with Applications to Function Approximation and Nonparametric Regression
by: Ou, Weigutian, et al.
Published: (2024)
by: Ou, Weigutian, et al.
Published: (2024)
Approximation Error and Complexity Bounds for ReLU Networks on Low-Regular Function Spaces
by: Davis, Owen, et al.
Published: (2024)
by: Davis, Owen, et al.
Published: (2024)
Quantitative Approximation Rates for Group Equivariant Learning
by: Siegel, Jonathan W., et al.
Published: (2026)
by: Siegel, Jonathan W., et al.
Published: (2026)
Simultaneous CNN Approximation on Manifolds with Applications to Boundary Value Problems
by: Zhou, Hanfei, et al.
Published: (2026)
by: Zhou, Hanfei, et al.
Published: (2026)
Optimal Approximation of Zonoids and Uniform Approximation by Shallow Neural Networks
by: Siegel, Jonathan W.
Published: (2023)
by: Siegel, Jonathan W.
Published: (2023)
Bayesian Neural Networks vs. Mixture Density Networks: Theoretical and Empirical Insights for Uncertainty-Aware Nonlinear Modeling
by: Ghosh, Riddhi Pratim, et al.
Published: (2025)
by: Ghosh, Riddhi Pratim, et al.
Published: (2025)
Efficient Approximation to Analytic and $L^p$ functions by Height-Augmented ReLU Networks
by: Li, ZeYu, et al.
Published: (2026)
by: Li, ZeYu, et al.
Published: (2026)
Shallow ReLU$^s$ Networks in $L^p$-Type and Sobolev Spaces: Approximation and Path-Norm Controlled Generalization
by: Li, Weizhao, et al.
Published: (2026)
by: Li, Weizhao, et al.
Published: (2026)
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
by: Gautam, Sushant, et al.
Published: (2026)
by: Gautam, Sushant, et al.
Published: (2026)
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs
by: Bahar, Atmane Ayoub Mansour, et al.
Published: (2024)
by: Bahar, Atmane Ayoub Mansour, et al.
Published: (2024)
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
On the existence of optimal shallow feedforward networks with ReLU activation
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
Error Analysis of Three-Layer Neural Network Trained with PGD for Deep Ritz Method
by: Jiao, Yuling, et al.
Published: (2024)
by: Jiao, Yuling, et al.
Published: (2024)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
by: Sarkar, Nilesh, et al.
Published: (2026)
by: Sarkar, Nilesh, et al.
Published: (2026)
Differentiable Neural Networks with RePU Activation: with Applications to Score Estimation and Isotonic Regression
by: Shen, Guohao, et al.
Published: (2023)
by: Shen, Guohao, et al.
Published: (2023)
Time-Frequency Analysis for Neural Networks
by: Abdeljawad, Ahmed, et al.
Published: (2025)
by: Abdeljawad, Ahmed, et al.
Published: (2025)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
by: Štefánik, Michal, et al.
Published: (2025)
by: Štefánik, Michal, et al.
Published: (2025)
ReLU neural network approximation to piecewise constant functions
by: Cai, Zhiqiang, et al.
Published: (2024)
by: Cai, Zhiqiang, et al.
Published: (2024)
Ridge Kernel Averaging and Uniform Approximation
by: Tian, James
Published: (2025)
by: Tian, James
Published: (2025)
Neural Networks Trained by Weight Permutation are Universal Approximators
by: Cai, Yongqiang, et al.
Published: (2024)
by: Cai, Yongqiang, et al.
Published: (2024)
A Hybrid CNN-Cheby-KAN Framework for Efficient Prediction of Two-Dimensional Airfoil Pressure Distribution
by: Chen, Yaohong, et al.
Published: (2025)
by: Chen, Yaohong, et al.
Published: (2025)
Large Language Models Report Subjective Experience Under Self-Referential Processing
by: Berg, Cameron, et al.
Published: (2025)
by: Berg, Cameron, et al.
Published: (2025)
Stealth edits to large language models
by: Sutton, Oliver J., et al.
Published: (2024)
by: Sutton, Oliver J., et al.
Published: (2024)
Equidistribution-based training of Free Knot Splines and ReLU Neural Networks
by: Appella, Simone, et al.
Published: (2024)
by: Appella, Simone, et al.
Published: (2024)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
Pay Attention to What You Need
by: Gao, Yifei, et al.
Published: (2023)
by: Gao, Yifei, et al.
Published: (2023)
A Hybrid Deep Learning and Anomaly Detection Framework for Real-Time Malicious URL Classification
by: Khaled, Berkani, et al.
Published: (2025)
by: Khaled, Berkani, et al.
Published: (2025)
On best approximation by multivariate ridge functions with applications to generalized translation networks
by: Geuchen, Paul, et al.
Published: (2024)
by: Geuchen, Paul, et al.
Published: (2024)
Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
by: Mitchell, Rupert, et al.
Published: (2025)
by: Mitchell, Rupert, et al.
Published: (2025)
Hyperplane Arrangements and Fixed Points in Iterated PWL Neural Networks
by: Beise, Hans-Peter
Published: (2024)
by: Beise, Hans-Peter
Published: (2024)
Discretization Error of Fourier Neural Operators
by: Lanthaler, Samuel, et al.
Published: (2024)
by: Lanthaler, Samuel, et al.
Published: (2024)
Standardized Threat Taxonomy for AI Security, Governance, and Regulatory Compliance
by: Huwyler, Hernan
Published: (2025)
by: Huwyler, Hernan
Published: (2025)
Time-Series Forecasting in Safety-Critical Environments: An EU-AI-Act-Compliant Open-Source Package / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein KI-VO-konformes Open-Source-Paket
by: Bartz-Beielstein, Thomas, et al.
Published: (2026)
by: Bartz-Beielstein, Thomas, et al.
Published: (2026)
Forging GEMs: Advancing Greek NLP through Quality-Based Corpus Curation
by: Apostolopoulou, Alexandra, et al.
Published: (2025)
by: Apostolopoulou, Alexandra, et al.
Published: (2025)
Universal Approximation of Dynamical Systems by Semi-Autonomous Neural ODEs and Applications
by: Li, Ziqian, et al.
Published: (2024)
by: Li, Ziqian, et al.
Published: (2024)
Mixed memories in Hopfield networks
by: Gayrard, Véronique
Published: (2025)
by: Gayrard, Véronique
Published: (2025)
MonoKAN: Certified Monotonic Kolmogorov-Arnold Network
by: Polo-Molina, Alejandro, et al.
Published: (2024)
by: Polo-Molina, Alejandro, et al.
Published: (2024)
Similar Items
-
Adaptive Randomized Neural Networks with Locally Activation Function: Theory and Algorithm for Solving PDEs
by: Bi, Ran, et al.
Published: (2026) -
Approximation bounds for norm constrained deep neural networks
by: Maiale, Francesco Paolo, et al.
Published: (2025) -
VORT: Adaptive Power-Law Memory for NLP Transformers
by: Mlaiki, Nabil
Published: (2026) -
Covering Numbers for Deep ReLU Networks with Applications to Function Approximation and Nonparametric Regression
by: Ou, Weigutian, et al.
Published: (2024) -
Approximation Error and Complexity Bounds for ReLU Networks on Low-Regular Function Spaces
by: Davis, Owen, et al.
Published: (2024)