EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Yau, Chung-Yiu, Li, Dawei, Glentis, Athanasios, Boreiko, Valentyn, Wai, Hoi-To, Hong, Mingyi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
A Stochastic Approximation Approach for Efficient Decentralized Optimization on Random Networks
by: Yau, Chung-Yiu, et al.
Published: (2024)
by: Yau, Chung-Yiu, et al.
Published: (2024)
EMC$^2$: Efficient MCMC Negative Sampling for Contrastive Learning with Global Convergence
by: Yau, Chung-Yiu, et al.
Published: (2024)
by: Yau, Chung-Yiu, et al.
Published: (2024)
Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates
by: Glentis, Athanasios, et al.
Published: (2026)
by: Glentis, Athanasios, et al.
Published: (2026)
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
by: Xie, Xingyu, et al.
Published: (2022)
by: Xie, Xingyu, et al.
Published: (2022)
Nesterov Acceleration for Ensemble Kalman Inversion and Variants
by: Vernon, Sydney, et al.
Published: (2025)
by: Vernon, Sydney, et al.
Published: (2025)
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023)
by: Liao, Fangshuo, et al.
Published: (2023)
Generalized Continuous-Time Models for Nesterov's Accelerated Gradient Methods
by: Park, Chanwoong, et al.
Published: (2024)
by: Park, Chanwoong, et al.
Published: (2024)
Inference of Online Newton Methods with Nesterov's Accelerated Sketching
by: Wang, Haoxuan, et al.
Published: (2026)
by: Wang, Haoxuan, et al.
Published: (2026)
A Concise Lyapunov Analysis of Nesterov's Accelerated Gradient Method
by: Liu, Jun
Published: (2025)
by: Liu, Jun
Published: (2025)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
by: Zimin, Aleksandr, et al.
Published: (2026)
by: Zimin, Aleksandr, et al.
Published: (2026)
Provable Acceleration of Nesterov's Accelerated Gradient for Rectangular Matrix Factorization and Linear Neural Networks
by: Xu, Zhenghao, et al.
Published: (2024)
by: Xu, Zhenghao, et al.
Published: (2024)
Nesterov acceleration in benignly non-convex landscapes
by: Gupta, Kanan, et al.
Published: (2024)
by: Gupta, Kanan, et al.
Published: (2024)
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
by: Borodich, Ekaterina, et al.
Published: (2025)
by: Borodich, Ekaterina, et al.
Published: (2025)
A Two-timescale Primal-dual Algorithm for Decentralized Optimization with Compression
by: Liu, Haoming, et al.
Published: (2025)
by: Liu, Haoming, et al.
Published: (2025)
Nesterov acceleration despite very noisy gradients
by: Gupta, Kanan, et al.
Published: (2023)
by: Gupta, Kanan, et al.
Published: (2023)
Continuized Nesterov Acceleration for Non-Convex Optimization
by: Hermant, Julien, et al.
Published: (2025)
by: Hermant, Julien, et al.
Published: (2025)
Nesterov Acceleration with Operator Decomposition
by: Lee, Jaewook, et al.
Published: (2026)
by: Lee, Jaewook, et al.
Published: (2026)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
by: Choudhury, Sayantan, et al.
Published: (2026)
by: Choudhury, Sayantan, et al.
Published: (2026)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022)
by: Liu, Xin, et al.
Published: (2022)
Nesterov Accelerated Distributed Optimization with Efficient Quantized Communication
by: Wu, Ruochen, et al.
Published: (2026)
by: Wu, Ruochen, et al.
Published: (2026)
Randomized Subspace Nesterov Accelerated Gradient
by: Omiya, Gaku, et al.
Published: (2026)
by: Omiya, Gaku, et al.
Published: (2026)
Decentralized Stochastic Optimization over Unreliable Networks via Two-timescales Updates
by: Liu, Haoming, et al.
Published: (2025)
by: Liu, Haoming, et al.
Published: (2025)
An Adaptive and Parameter-Free Nesterov's Accelerated Gradient Method for Convex Optimization
by: Suh, Jaewook J., et al.
Published: (2025)
by: Suh, Jaewook J., et al.
Published: (2025)
Heavy Ball and Nesterov Accelerations with Hessian-driven Damping for Nonconvex Optimization
by: Hadjisavvas, N., et al.
Published: (2025)
by: Hadjisavvas, N., et al.
Published: (2025)
The Iterates of Nesterov's Accelerated Algorithm Converge in The Critical Regimes
by: Bot, Radu Ioan, et al.
Published: (2025)
by: Bot, Radu Ioan, et al.
Published: (2025)
A Nesterov-Accelerated Primal-Dual Splitting Algorithm for Convex Nonsmooth Optimization
by: Condat, Laurent, et al.
Published: (2026)
by: Condat, Laurent, et al.
Published: (2026)
Accelerated linearized alternating direction method of multipliers with Nesterov extrapolation
by: He, X., et al.
Published: (2023)
by: He, X., et al.
Published: (2023)
Technical Report: A Totally Asynchronous Nesterov's Accelerated Gradient Method for Convex Optimization
by: Pond, Ellie, et al.
Published: (2024)
by: Pond, Ellie, et al.
Published: (2024)
Point Convergence of Nesterov's Accelerated Gradient Method: An AI-Assisted Proof
by: Jang, Uijeong, et al.
Published: (2025)
by: Jang, Uijeong, et al.
Published: (2025)
Fractional-Order Nesterov Dynamics for Convex Optimization
by: Ranoto, Tumelo
Published: (2025)
by: Ranoto, Tumelo
Published: (2025)
Asynchronous and Stochastic Distributed Resource Allocation
by: Li, Qiang, et al.
Published: (2025)
by: Li, Qiang, et al.
Published: (2025)
A Nesterov-style Accelerated Gradient Descent Algorithm for the Symmetric Eigenvalue Problem
by: Alimisis, Foivos, et al.
Published: (2024)
by: Alimisis, Foivos, et al.
Published: (2024)
The Nesterov-Spokoiny Acceleration Achieves Strict $o(1/k^2)$ Convergence
by: Peng, Weibin, et al.
Published: (2023)
by: Peng, Weibin, et al.
Published: (2023)
EARL-BO: Reinforcement Learning for Multi-Step Lookahead, High-Dimensional Bayesian Optimization
by: Cheon, Mujin, et al.
Published: (2024)
by: Cheon, Mujin, et al.
Published: (2024)
Omega: Optimistic EMA Gradients
by: Ramirez, Juan, et al.
Published: (2023)
by: Ramirez, Juan, et al.
Published: (2023)
From Halpern's Fixed-Point Iterations to Nesterov's Accelerated Interpretations for Root-Finding Problems
by: Tran-Dinh, Quoc
Published: (2022)
by: Tran-Dinh, Quoc
Published: (2022)
Stochastic Gradient Descent with Strategic Querying
by: Jiang, Nanfei, et al.
Published: (2025)
by: Jiang, Nanfei, et al.
Published: (2025)
Understanding Lookahead Dynamics Through Laplace Transform
by: Sanyal, Aniket, et al.
Published: (2025)
by: Sanyal, Aniket, et al.
Published: (2025)
Delayed supermartingale convergence lemmas for stochastic approximation with Nesterov momentum
by: Ming-Kun, Zhang
Published: (2024)
by: Ming-Kun, Zhang
Published: (2024)
Similar Items
-
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025) -
A Stochastic Approximation Approach for Efficient Decentralized Optimization on Random Networks
by: Yau, Chung-Yiu, et al.
Published: (2024) -
EMC$^2$: Efficient MCMC Negative Sampling for Contrastive Learning with Global Convergence
by: Yau, Chung-Yiu, et al.
Published: (2024) -
Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates
by: Glentis, Athanasios, et al.
Published: (2026) -
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
by: Xie, Xingyu, et al.
Published: (2022)