Unraveling the Gradient Descent Dynamics of Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Bingqing, Han, Boran, Zhang, Shuai, Ding, Jie, Hong, Mingyi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
by: Qin, Zhen, et al.
Published: (2025)
by: Qin, Zhen, et al.
Published: (2025)
On Penalty-based Bilevel Gradient Descent Method
by: Shen, Han, et al.
Published: (2023)
by: Shen, Han, et al.
Published: (2023)
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
by: Ozkara, Kaan, et al.
Published: (2024)
by: Ozkara, Kaan, et al.
Published: (2024)
Stochastic Adaptive Gradient Descent Without Descent
by: Aujol, Jean-François, et al.
Published: (2025)
by: Aujol, Jean-François, et al.
Published: (2025)
Anytime Acceleration of Gradient Descent
by: Zhang, Zihan, et al.
Published: (2024)
by: Zhang, Zihan, et al.
Published: (2024)
Corner Gradient Descent
by: Yarotsky, Dmitry
Published: (2025)
by: Yarotsky, Dmitry
Published: (2025)
Parameter-free Clipped Gradient Descent Meets Polyak
by: Takezawa, Yuki, et al.
Published: (2024)
by: Takezawa, Yuki, et al.
Published: (2024)
Adaptive Conditional Gradient Descent
by: Khademi, Abbas, et al.
Published: (2025)
by: Khademi, Abbas, et al.
Published: (2025)
$k$-SVD with Gradient Descent
by: Jedra, Yassir, et al.
Published: (2025)
by: Jedra, Yassir, et al.
Published: (2025)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
Stochastic Gradient Descent with Adaptive Data
by: Che, Ethan, et al.
Published: (2024)
by: Che, Ethan, et al.
Published: (2024)
Stochastic Gradient Descent with Strategic Querying
by: Jiang, Nanfei, et al.
Published: (2025)
by: Jiang, Nanfei, et al.
Published: (2025)
Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix Factorization
by: Zhang, Gavin, et al.
Published: (2025)
by: Zhang, Gavin, et al.
Published: (2025)
Fast and Accurate Estimation of Low-Rank Matrices from Noisy Measurements via Preconditioned Non-Convex Gradient Descent
by: Zhang, Gavin, et al.
Published: (2023)
by: Zhang, Gavin, et al.
Published: (2023)
Low-Tubal-Rank Tensor Recovery via Factorized Gradient Descent
by: Liu, Zhiyu, et al.
Published: (2024)
by: Liu, Zhiyu, et al.
Published: (2024)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
by: Arda, Enes, et al.
Published: (2026)
by: Arda, Enes, et al.
Published: (2026)
On the Convergence of Stochastic Gradient Descent with Perturbed Forward-Backward Passes
by: Kong, Boao, et al.
Published: (2026)
by: Kong, Boao, et al.
Published: (2026)
Mirror and Preconditioned Gradient Descent in Wasserstein Space
by: Bonet, Clément, et al.
Published: (2024)
by: Bonet, Clément, et al.
Published: (2024)
Derivatives of Stochastic Gradient Descent in parametric optimization
by: Iutzeler, Franck, et al.
Published: (2024)
by: Iutzeler, Franck, et al.
Published: (2024)
Enhancing Fractional Gradient Descent with Learned Optimizers
by: Sobotka, Jan, et al.
Published: (2025)
by: Sobotka, Jan, et al.
Published: (2025)
Convergence of Alternating Gradient Descent for Matrix Factorization
by: Ward, Rachel, et al.
Published: (2023)
by: Ward, Rachel, et al.
Published: (2023)
A Local Polyak-Lojasiewicz and Descent Lemma of Gradient Descent For Overparametrized Linear Models
by: Xu, Ziqing, et al.
Published: (2025)
by: Xu, Ziqing, et al.
Published: (2025)
Efficient Low-Tubal-Rank Tensor Estimation via Alternating Preconditioned Gradient Descent
by: Liu, Zhiyu, et al.
Published: (2025)
by: Liu, Zhiyu, et al.
Published: (2025)
Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's Law
by: Kunstner, Frederik, et al.
Published: (2025)
by: Kunstner, Frederik, et al.
Published: (2025)
Almost Bayesian: The Fractal Dynamics of Stochastic Gradient Descent
by: Hennick, Max, et al.
Published: (2025)
by: Hennick, Max, et al.
Published: (2025)
The Sample Complexity of Gradient Descent in Stochastic Convex Optimization
by: Livni, Roi
Published: (2024)
by: Livni, Roi
Published: (2024)
Open Problem: Anytime Convergence Rate of Gradient Descent
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2024)
by: Ziyin, Liu, et al.
Published: (2024)
Convergence Analysis of Stochastic Gradient Descent with MCMC Estimators
by: Li, Tianyou, et al.
Published: (2023)
by: Li, Tianyou, et al.
Published: (2023)
On Gradient Descent Ascent for Nonconvex-Concave Minimax Problems
by: Lin, Tianyi, et al.
Published: (2019)
by: Lin, Tianyi, et al.
Published: (2019)
Gradient Descent's Last Iterate is Often (slightly) Suboptimal
by: Kornowski, Guy, et al.
Published: (2026)
by: Kornowski, Guy, et al.
Published: (2026)
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
by: Köhne, Frederik, et al.
Published: (2023)
by: Köhne, Frederik, et al.
Published: (2023)
Gauss-Newton Natural Gradient Descent for Shape Learning
by: King, James, et al.
Published: (2026)
by: King, James, et al.
Published: (2026)
Functional Central Limit Theorem for Stochastic Gradient Descent
by: Flamand, Kessang, et al.
Published: (2026)
by: Flamand, Kessang, et al.
Published: (2026)
On the Inherent Privacy of Zeroth Order Projected Gradient Descent
by: Gupta, Devansh, et al.
Published: (2025)
by: Gupta, Devansh, et al.
Published: (2025)
Preconditioned Gradient Descent for Overparameterized Nonconvex Burer--Monteiro Factorization with Global Optimality Certification
by: Zhang, Gavin, et al.
Published: (2022)
by: Zhang, Gavin, et al.
Published: (2022)
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
by: Phunyaphibarn, Prin, et al.
Published: (2023)
by: Phunyaphibarn, Prin, et al.
Published: (2023)
Quadratic Gradient: A Unified Framework Bridging Gradient Descent and Newton-Type Methods by Synthesizing Hessians and Gradients
by: Chiang, John
Published: (2022)
by: Chiang, John
Published: (2022)
Similar Items
-
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025) -
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
by: Qin, Zhen, et al.
Published: (2025) -
On Penalty-based Bilevel Gradient Descent Method
by: Shen, Han, et al.
Published: (2023) -
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
by: Ozkara, Kaan, et al.
Published: (2024) -
Stochastic Adaptive Gradient Descent Without Descent
by: Aujol, Jean-François, et al.
Published: (2025)