A unified framework for establishing the universal approximation of transformer-type architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Jingpu, Lin, Ting, Shen, Zuowei, Li, Qianxiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep learning and the rate of approximation by flows
by: Cheng, Jingpu, et al.
Published: (2026)
by: Cheng, Jingpu, et al.
Published: (2026)
Unifying back-propagation and forward-forward algorithms through model predictive control
by: Ren, Lianhai, et al.
Published: (2024)
by: Ren, Lianhai, et al.
Published: (2024)
One to beat them all: "RYU" -- a unifying framework for the construction of safe balls
by: Tran, Thu-Le, et al.
Published: (2023)
by: Tran, Thu-Le, et al.
Published: (2023)
A unified perspective on fine-tuning and sampling with diffusion and flow models
by: Domingo-Enrich, Carles, et al.
Published: (2026)
by: Domingo-Enrich, Carles, et al.
Published: (2026)
Beyond likelihood ratio bias: Nested multi-time-scale stochastic approximation for likelihood-free parameter estimation
by: Li, Zehao, et al.
Published: (2024)
by: Li, Zehao, et al.
Published: (2024)
Surrogate-based optimization of system architectures subject to hidden constraints
by: Bussemaker, Jasper, et al.
Published: (2025)
by: Bussemaker, Jasper, et al.
Published: (2025)
Global optimization of graph acquisition functions for neural architecture search
by: Xie, Yilin, et al.
Published: (2025)
by: Xie, Yilin, et al.
Published: (2025)
Tightening convex relaxations of trained neural networks: a unified approach for convex and S-shaped activations
by: Carrasco, Pablo, et al.
Published: (2024)
by: Carrasco, Pablo, et al.
Published: (2024)
Randomized multi-class classification under system constraints: a unified approach via post-processing
by: Chzhen, Evgenii, et al.
Published: (2025)
by: Chzhen, Evgenii, et al.
Published: (2025)
Exploiting weight-space symmetries for approximating curvature
by: Artemev, Artem, et al.
Published: (2026)
by: Artemev, Artem, et al.
Published: (2026)
Constructive approximate transport maps with normalizing flows
by: Álvarez-López, Antonio, et al.
Published: (2024)
by: Álvarez-López, Antonio, et al.
Published: (2024)
Learning based convex approximation for constrained parametric optimization
by: Liu, Kang, et al.
Published: (2025)
by: Liu, Kang, et al.
Published: (2025)
Size and depth of monotone neural networks: interpolation and approximation
by: Mikulincer, Dan, et al.
Published: (2022)
by: Mikulincer, Dan, et al.
Published: (2022)
A mathematical framework for time-delay reservoir computing analysis
by: Clabaut, Anh-Tuan, et al.
Published: (2026)
by: Clabaut, Anh-Tuan, et al.
Published: (2026)
Accelerated stochastic approximation with state-dependent noise
by: Ilandarideva, Sasila, et al.
Published: (2023)
by: Ilandarideva, Sasila, et al.
Published: (2023)
Terminally constrained flow-based generative models from an optimal control perspective
by: Gao, Weiguo, et al.
Published: (2026)
by: Gao, Weiguo, et al.
Published: (2026)
Automatic nonlinear MPC approximation with closed-loop guarantees
by: Tokmak, Abdullah, et al.
Published: (2023)
by: Tokmak, Abdullah, et al.
Published: (2023)
Model approximation in MDPs with unbounded per-step cost
by: Bozkurt, Berk, et al.
Published: (2024)
by: Bozkurt, Berk, et al.
Published: (2024)
PACSBO: Probably approximately correct safe Bayesian optimization
by: Tokmak, Abdullah, et al.
Published: (2024)
by: Tokmak, Abdullah, et al.
Published: (2024)
Provable optimal transport with transformers: The essence of depth and prompt engineering
by: Daneshmand, Hadi
Published: (2024)
by: Daneshmand, Hadi
Published: (2024)
A framework for bilevel optimization that enables stochastic and global variance reduction algorithms
by: Dagréou, Mathieu, et al.
Published: (2022)
by: Dagréou, Mathieu, et al.
Published: (2022)
A successive approximation method in functional spaces for hierarchical optimal control problems and its application to learning
by: Befekadu, Getachew K.
Published: (2024)
by: Befekadu, Getachew K.
Published: (2024)
LoDAdaC: a unified local training-based decentralized framework with adaptive gradients and compressed communication
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Robust Neural IDA-PBC: passivity-based stabilization under approximations
by: Sanchez-Escalonilla, Santiago, et al.
Published: (2024)
by: Sanchez-Escalonilla, Santiago, et al.
Published: (2024)
Penalty-Based First-Order Methods for Bilevel Optimization with Minimax and Constrained Lower-Level Problems
by: Shen, Yiyang, et al.
Published: (2026)
by: Shen, Yiyang, et al.
Published: (2026)
A Sinkhorn-type Algorithm for Constrained Optimal Transport
by: Tang, Xun, et al.
Published: (2024)
by: Tang, Xun, et al.
Published: (2024)
Robust stabilization of polytopic systems via fast and reliable neural network-based approximations
by: Fabiani, Filippo, et al.
Published: (2022)
by: Fabiani, Filippo, et al.
Published: (2022)
A Hessian-Aware Stochastic Differential Equation for Modelling SGD
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
A block-coordinate descent framework for non-convex composite optimization. Application to sparse precision matrix estimation
by: Lauga, Guillaume
Published: (2026)
by: Lauga, Guillaume
Published: (2026)
Bayesian optimization as a flexible and efficient design framework for sustainable process systems
by: Paulson, Joel A., et al.
Published: (2024)
by: Paulson, Joel A., et al.
Published: (2024)
SafEDMD: A Koopman-based data-driven controller design framework for nonlinear dynamical systems
by: Strässer, Robin, et al.
Published: (2024)
by: Strässer, Robin, et al.
Published: (2024)
A stochastic smoothing framework for nonconvex-nonconcave min-sum-max problems with applications to Wasserstein distributionally robust optimization
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Instance-optimal stochastic convex optimization: Can we improve upon sample-average and robust stochastic approximation?
by: Jiang, Liwei, et al.
Published: (2026)
by: Jiang, Liwei, et al.
Published: (2026)
Stochastic Hessian Fittings with Lie Groups
by: Li, Xi-Lin
Published: (2024)
by: Li, Xi-Lin
Published: (2024)
The Convergence of Dynamic Routing between Capsules
by: Ye, Daoyuan, et al.
Published: (2025)
by: Ye, Daoyuan, et al.
Published: (2025)
Towards Understanding Generalization and Stability Gaps between Centralized and Decentralized Federated Learning
by: Sun, Yan, et al.
Published: (2023)
by: Sun, Yan, et al.
Published: (2023)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
ODE approximation for the Adam algorithm: General and overparametrized setting
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Online estimation of the inverse of the Hessian for stochastic optimization with application to universal stochastic Newton algorithms
by: Godichon-Baggioni, Antoine, et al.
Published: (2024)
by: Godichon-Baggioni, Antoine, et al.
Published: (2024)
General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Similar Items
-
Deep learning and the rate of approximation by flows
by: Cheng, Jingpu, et al.
Published: (2026) -
Unifying back-propagation and forward-forward algorithms through model predictive control
by: Ren, Lianhai, et al.
Published: (2024) -
One to beat them all: "RYU" -- a unifying framework for the construction of safe balls
by: Tran, Thu-Le, et al.
Published: (2023) -
A unified perspective on fine-tuning and sampling with diffusion and flow models
by: Domingo-Enrich, Carles, et al.
Published: (2026) -
Beyond likelihood ratio bias: Nested multi-time-scale stochastic approximation for likelihood-free parameter estimation
by: Li, Zehao, et al.
Published: (2024)