Provable Acceleration of Nesterov's Accelerated Gradient for Rectangular Matrix Factorization and Linear Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Zhenghao, Wang, Yuqing, Zhao, Tuo, Ward, Rachel, Tao, Molei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022)
by: Liu, Xin, et al.
Published: (2022)
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023)
by: Liao, Fangshuo, et al.
Published: (2023)
Generalized Continuous-Time Models for Nesterov's Accelerated Gradient Methods
by: Park, Chanwoong, et al.
Published: (2024)
by: Park, Chanwoong, et al.
Published: (2024)
A Concise Lyapunov Analysis of Nesterov's Accelerated Gradient Method
by: Liu, Jun
Published: (2025)
by: Liu, Jun
Published: (2025)
Convergence of Alternating Gradient Descent for Matrix Factorization
by: Ward, Rachel, et al.
Published: (2023)
by: Ward, Rachel, et al.
Published: (2023)
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
by: Yau, Chung-Yiu, et al.
Published: (2026)
by: Yau, Chung-Yiu, et al.
Published: (2026)
Inference of Online Newton Methods with Nesterov's Accelerated Sketching
by: Wang, Haoxuan, et al.
Published: (2026)
by: Wang, Haoxuan, et al.
Published: (2026)
Nesterov Acceleration for Ensemble Kalman Inversion and Variants
by: Vernon, Sydney, et al.
Published: (2025)
by: Vernon, Sydney, et al.
Published: (2025)
How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
by: Lai, Kuo-Wei, et al.
Published: (2026)
by: Lai, Kuo-Wei, et al.
Published: (2026)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
by: Zimin, Aleksandr, et al.
Published: (2026)
by: Zimin, Aleksandr, et al.
Published: (2026)
Evaluating the design space of diffusion-based generative models
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
Regularized Gradient Clipping Provably Trains Wide and Deep Neural Networks
by: Tucat, Matteo, et al.
Published: (2024)
by: Tucat, Matteo, et al.
Published: (2024)
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
by: Borodich, Ekaterina, et al.
Published: (2025)
by: Borodich, Ekaterina, et al.
Published: (2025)
Anytime Acceleration of Gradient Descent
by: Zhang, Zihan, et al.
Published: (2024)
by: Zhang, Zihan, et al.
Published: (2024)
Gradient-Variation Online Adaptivity for Accelerated Optimization with Hölder Smoothness
by: Zhao, Yuheng, et al.
Published: (2025)
by: Zhao, Yuheng, et al.
Published: (2025)
Randomized Subspace Nesterov Accelerated Gradient
by: Omiya, Gaku, et al.
Published: (2026)
by: Omiya, Gaku, et al.
Published: (2026)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
Towards Simple and Provable Parameter-Free Adaptive Gradient Methods
by: Tao, Yuanzhe, et al.
Published: (2024)
by: Tao, Yuanzhe, et al.
Published: (2024)
Provable Acceleration for Diffusion Models under Minimal Assumptions
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
by: Jacot, Arthur, et al.
Published: (2024)
by: Jacot, Arthur, et al.
Published: (2024)
Revisiting Gradient Normalization and Clipping for Nonconvex SGD under Heavy-Tailed Noise: Necessity, Sufficiency, and Acceleration
by: Sun, Tao, et al.
Published: (2024)
by: Sun, Tao, et al.
Published: (2024)
Accelerated Optimization Landscape of Linear-Quadratic Regulator
by: Feng, Lechen, et al.
Published: (2023)
by: Feng, Lechen, et al.
Published: (2023)
Accelerating Single-Pass SGD for Generalized Linear Prediction
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Nesterov Acceleration with Operator Decomposition
by: Lee, Jaewook, et al.
Published: (2026)
by: Lee, Jaewook, et al.
Published: (2026)
Matrix Completion with Graph Information: A Provable Nonconvex Optimization Approach
by: Wang, Yao, et al.
Published: (2025)
by: Wang, Yao, et al.
Published: (2025)
Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient Noise
by: Pan, Rui, et al.
Published: (2023)
by: Pan, Rui, et al.
Published: (2023)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
An Accelerated Gradient Method for Convex Smooth Simple Bilevel Optimization
by: Cao, Jincheng, et al.
Published: (2024)
by: Cao, Jincheng, et al.
Published: (2024)
Quantitative Convergences of Lie Group Momentum Optimizers
by: Kong, Lingkai, et al.
Published: (2024)
by: Kong, Lingkai, et al.
Published: (2024)
CRONOS: Enhancing Deep Learning with Scalable GPU Accelerated Convex Neural Networks
by: Feng, Miria, et al.
Published: (2024)
by: Feng, Miria, et al.
Published: (2024)
PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network Training
by: Yang, Shenghao, et al.
Published: (2026)
by: Yang, Shenghao, et al.
Published: (2026)
Nesterov acceleration in benignly non-convex landscapes
by: Gupta, Kanan, et al.
Published: (2024)
by: Gupta, Kanan, et al.
Published: (2024)
An Adaptive and Parameter-Free Nesterov's Accelerated Gradient Method for Convex Optimization
by: Suh, Jaewook J., et al.
Published: (2025)
by: Suh, Jaewook J., et al.
Published: (2025)
Point Convergence of Nesterov's Accelerated Gradient Method: An AI-Assisted Proof
by: Jang, Uijeong, et al.
Published: (2025)
by: Jang, Uijeong, et al.
Published: (2025)
Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix Factorization
by: Zhang, Gavin, et al.
Published: (2025)
by: Zhang, Gavin, et al.
Published: (2025)
Stochastic Constrained Decentralized Optimization for Machine Learning with Fewer Data Oracles: a Gradient Sliding Approach
by: Nguyen, Hoang Huy, et al.
Published: (2024)
by: Nguyen, Hoang Huy, et al.
Published: (2024)
Provably-Stable Neural Network-Based Control of Nonlinear Systems
by: Li, Anran, et al.
Published: (2025)
by: Li, Anran, et al.
Published: (2025)
Accelerated Gradient Tracking over Time-varying Graphs for Decentralized Optimization
by: Li, Huan, et al.
Published: (2021)
by: Li, Huan, et al.
Published: (2021)
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
by: Vaswani, Sharan, et al.
Published: (2021)
by: Vaswani, Sharan, et al.
Published: (2021)
From Inexact Gradients to Byzantine Robustness: Acceleration and Optimization under Similarity
by: Gaucher, Renaud, et al.
Published: (2026)
by: Gaucher, Renaud, et al.
Published: (2026)
Similar Items
-
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022) -
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023) -
Generalized Continuous-Time Models for Nesterov's Accelerated Gradient Methods
by: Park, Chanwoong, et al.
Published: (2024) -
A Concise Lyapunov Analysis of Nesterov's Accelerated Gradient Method
by: Liu, Jun
Published: (2025) -
Convergence of Alternating Gradient Descent for Matrix Factorization
by: Ward, Rachel, et al.
Published: (2023)