How Transformers Get Rich: Approximation and Dynamics Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Mingze, Yu, Ruoxi, E, Weinan, Wu, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
GradPower: Powering Gradients for Faster Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
Achieving Margin Maximization Exponentially Fast via Progressive Norm Rescaling
by: Wang, Mingze, et al.
Published: (2023)
by: Wang, Mingze, et al.
Published: (2023)
Improving Generalization and Convergence by Enhancing Implicit Regularization
by: Wang, Mingze, et al.
Published: (2024)
by: Wang, Mingze, et al.
Published: (2024)
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2024)
by: Ziyin, Liu, et al.
Published: (2024)
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
by: Wang, Jinbo, et al.
Published: (2026)
by: Wang, Jinbo, et al.
Published: (2026)
A Theoretical Analysis of Self-Supervised Learning for Vision Transformers
by: Huang, Yu, et al.
Published: (2024)
by: Huang, Yu, et al.
Published: (2024)
Unraveling the Gradient Descent Dynamics of Transformers
by: Song, Bingqing, et al.
Published: (2024)
by: Song, Bingqing, et al.
Published: (2024)
Understanding Lookahead Dynamics Through Laplace Transform
by: Sanyal, Aniket, et al.
Published: (2025)
by: Sanyal, Aniket, et al.
Published: (2025)
NOMADS: Non-Markovian Optimization-based Modeling for Approximate Dynamics with Spatially-homogeneous Memory
by: Anzaki, Ryoji, et al.
Published: (2024)
by: Anzaki, Ryoji, et al.
Published: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Enhancing Distributional Robustness in Principal Component Analysis by Wasserstein Distances
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
A Simple Finite-Time Analysis of TD Learning with Linear Function Approximation
by: Mitra, Aritra
Published: (2024)
by: Mitra, Aritra
Published: (2024)
Gaussian Approximation and Multiplier Bootstrap for Federated Linear Stochastic Approximation
by: Levin, Ilya, et al.
Published: (2026)
by: Levin, Ilya, et al.
Published: (2026)
Follow The Approximate Sparse Leader for No-Regret Online Sparse Linear Approximation
by: Mukhopadhyay, Samrat, et al.
Published: (2025)
by: Mukhopadhyay, Samrat, et al.
Published: (2025)
Heavy-Tailed and Long-Range Dependent Noise in Stochastic Approximation: A Finite-Time Analysis
by: Chandak, Siddharth, et al.
Published: (2026)
by: Chandak, Siddharth, et al.
Published: (2026)
A Finite-Time Analysis of TD Learning with Linear Function Approximation without Projections or Strong Convexity
by: Lee, Wei-Cheng, et al.
Published: (2025)
by: Lee, Wei-Cheng, et al.
Published: (2025)
Central Limit Theorem for Two-Time-Scale Approximate Distributionally Robust RL
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
Tame Riemannian Stochastic Approximation
by: Aspman, Johannes, et al.
Published: (2023)
by: Aspman, Johannes, et al.
Published: (2023)
Blackwell's Approachability with Approximation Algorithms
by: Garber, Dan, et al.
Published: (2025)
by: Garber, Dan, et al.
Published: (2025)
Exponential Concentration in Stochastic Approximation
by: Law, Kody, et al.
Published: (2022)
by: Law, Kody, et al.
Published: (2022)
Reinforcement Learning from Partial Observation: Linear Function Approximation with Provable Sample Efficiency
by: Cai, Qi, et al.
Published: (2022)
by: Cai, Qi, et al.
Published: (2022)
On the Global Convergence of Risk-Averse Natural Policy Gradient Methods with Expected Conditional Risk Measures
by: Yu, Xian, et al.
Published: (2023)
by: Yu, Xian, et al.
Published: (2023)
Federated Temporal Difference Learning with Linear Function Approximation under Environmental Heterogeneity
by: Wang, Han, et al.
Published: (2023)
by: Wang, Han, et al.
Published: (2023)
How Well Can Transformers Emulate In-context Newton's Method?
by: Giannou, Angeliki, et al.
Published: (2024)
by: Giannou, Angeliki, et al.
Published: (2024)
Stochastic Approximation with Block Coordinate Optimal Stepsizes
by: Jiang, Tao, et al.
Published: (2025)
by: Jiang, Tao, et al.
Published: (2025)
Online Covariance Estimation in Nonsmooth Stochastic Approximation
by: Jiang, Liwei, et al.
Published: (2025)
by: Jiang, Liwei, et al.
Published: (2025)
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
by: Semenov, Andrei, et al.
Published: (2025)
by: Semenov, Andrei, et al.
Published: (2025)
Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-Critic
by: Zhang, Yufeng, et al.
Published: (2021)
by: Zhang, Yufeng, et al.
Published: (2021)
Random Features Approximation for Control-Affine Systems
by: Kazemian, Kimia, et al.
Published: (2024)
by: Kazemian, Kimia, et al.
Published: (2024)
Black-Box Approximation and Optimization with Hierarchical Tucker Decomposition
by: Ryzhakov, Gleb, et al.
Published: (2024)
by: Ryzhakov, Gleb, et al.
Published: (2024)
Weakly Time-Coupled Approximation of Markov Decision Processes
by: Soheili, Negar, et al.
Published: (2026)
by: Soheili, Negar, et al.
Published: (2026)
Turbocharging Gaussian Process Inference with Approximate Sketch-and-Project
by: Rathore, Pratik, et al.
Published: (2025)
by: Rathore, Pratik, et al.
Published: (2025)
Efficient Stochastic Approximation of Minimax Excess Risk Optimization
by: Zhang, Lijun, et al.
Published: (2023)
by: Zhang, Lijun, et al.
Published: (2023)
Scalable Approximate Algorithms for Optimal Transport Linear Models
by: Kacprzak, Tomasz, et al.
Published: (2025)
by: Kacprzak, Tomasz, et al.
Published: (2025)
Reinforcement Learning with Function Approximation for Non-Markov Processes
by: Kara, Ali Devran
Published: (2026)
by: Kara, Ali Devran
Published: (2026)
Stochastic Approximation Methods for Distortion Risk Measure Optimization
by: Jiang, Jinyang, et al.
Published: (2025)
by: Jiang, Jinyang, et al.
Published: (2025)
Tractable Representations for Convergent Approximation of Distributional HJB Equations
by: Alhosh, Julie, et al.
Published: (2025)
by: Alhosh, Julie, et al.
Published: (2025)
Concentration of Contractive Stochastic Approximation: Additive and Multiplicative Noise
by: Chen, Zaiwei, et al.
Published: (2023)
by: Chen, Zaiwei, et al.
Published: (2023)
A Retrospective Approximation Approach for Smooth Stochastic Optimization
by: Newton, David, et al.
Published: (2021)
by: Newton, David, et al.
Published: (2021)
Similar Items
-
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025) -
GradPower: Powering Gradients for Faster Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025) -
Achieving Margin Maximization Exponentially Fast via Progressive Norm Rescaling
by: Wang, Mingze, et al.
Published: (2023) -
Improving Generalization and Convergence by Enhancing Implicit Regularization
by: Wang, Mingze, et al.
Published: (2024) -
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2024)