Improved Analysis for Sign-based Methods with Momentum Updates
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Wei, Yu, Dingzhi, Yang, Sifan, Yang, Wenhao, Zhang, Lijun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Sign-Based Optimization: Accelerating Convergence via Variance Reduction
by: Jiang, Wei, et al.
Published: (2024)
by: Jiang, Wei, et al.
Published: (2024)
Projection-Free Variance Reduction Methods for Stochastic Constrained Multi-Level Compositional Optimization
by: Jiang, Wei, et al.
Published: (2024)
by: Jiang, Wei, et al.
Published: (2024)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
by: Tao, Hongyi, et al.
Published: (2026)
by: Tao, Hongyi, et al.
Published: (2026)
Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Adaptive Variance Reduction for Stochastic Optimization under Weaker Assumptions
by: Jiang, Wei, et al.
Published: (2024)
by: Jiang, Wei, et al.
Published: (2024)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Group Distributionally Robust Optimization with Flexible Sample Queries
by: Bai, Haomin, et al.
Published: (2025)
by: Bai, Haomin, et al.
Published: (2025)
Convergence Analysis of the Lion Optimizer in Centralized and Distributed Settings
by: Jiang, Wei, et al.
Published: (2025)
by: Jiang, Wei, et al.
Published: (2025)
Efficient Penalty-Based Bilevel Methods: Improved Analysis, Novel Updates, and Flatness Condition
by: Jiang, Liuyuan, et al.
Published: (2025)
by: Jiang, Liuyuan, et al.
Published: (2025)
Universal Online Convex Optimization with $1$ Projection per Round
by: Yang, Wenhao, et al.
Published: (2024)
by: Yang, Wenhao, et al.
Published: (2024)
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
by: Everett, Katie, et al.
Published: (2026)
by: Everett, Katie, et al.
Published: (2026)
Better LMO-based Momentum Methods with Second-Order Information
by: Khirirat, Sarit, et al.
Published: (2025)
by: Khirirat, Sarit, et al.
Published: (2025)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
Improving Stochastic Cubic Newton with Momentum
by: Chayti, El Mahdi, et al.
Published: (2024)
by: Chayti, El Mahdi, et al.
Published: (2024)
Compressed Decentralized Momentum Stochastic Gradient Methods for Nonconvex Optimization
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
An Iteratively Reweighted Method for Sparse Optimization on Nonconvex $\ell_{p}$ Ball
by: Wang, Hao, et al.
Published: (2021)
by: Wang, Hao, et al.
Published: (2021)
Efficient Stochastic Approximation of Minimax Excess Risk Optimization
by: Zhang, Lijun, et al.
Published: (2023)
by: Zhang, Lijun, et al.
Published: (2023)
Improved Convergence in Parameter-Agnostic Error Feedback through Momentum
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
Communication-Efficient Gradient Descent-Accent Methods for Distributed Variational Inequalities: Unified Analysis and Local Updates
by: Zhang, Siqi, et al.
Published: (2023)
by: Zhang, Siqi, et al.
Published: (2023)
Stochastic Gradient Methods with Preconditioned Updates
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025)
by: Kovalev, Dmitry
Published: (2025)
Neighbor-Sampling Based Momentum Stochastic Methods for Training Graph Neural Networks
by: Noel, Molly, et al.
Published: (2025)
by: Noel, Molly, et al.
Published: (2025)
First and Second Order Approximations to Stochastic Gradient Descent Methods with Momentum Terms
by: Lu, Eric
Published: (2025)
by: Lu, Eric
Published: (2025)
When is Momentum Extragradient Optimal? A Polynomial-Based Analysis
by: Kim, Junhyung Lyle, et al.
Published: (2022)
by: Kim, Junhyung Lyle, et al.
Published: (2022)
Towards Fully Parameter-Free Stochastic Optimization: Grid Search with Self-Bounding Analysis
by: Zhao, Yuheng, et al.
Published: (2026)
by: Zhao, Yuheng, et al.
Published: (2026)
Universal Online Convex Optimization Meets Second-order Bounds
by: Zhang, Lijun, et al.
Published: (2021)
by: Zhang, Lijun, et al.
Published: (2021)
Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF
by: Shen, Han, et al.
Published: (2024)
by: Shen, Han, et al.
Published: (2024)
Global Convergence of Natural Policy Gradient with Hessian-aided Momentum Variance Reduction
by: Feng, Jie, et al.
Published: (2024)
by: Feng, Jie, et al.
Published: (2024)
Projection-free Online Learning over Strongly Convex Sets
by: Wan, Yuanyu, et al.
Published: (2020)
by: Wan, Yuanyu, et al.
Published: (2020)
Random Scaling and Momentum for Non-smooth Non-convex Optimization
by: Zhang, Qinzi, et al.
Published: (2024)
by: Zhang, Qinzi, et al.
Published: (2024)
Elementary Analysis of Policy Gradient Methods
by: Liu, Jiacai, et al.
Published: (2024)
by: Liu, Jiacai, et al.
Published: (2024)
Continuous-time q-Learning for Jump-Diffusion Models under Tsallis Entropy
by: Bo, Lijun, et al.
Published: (2024)
by: Bo, Lijun, et al.
Published: (2024)
Adaptive Delayed-Update Cyclic Algorithm for Variational Inequalities
by: Wei, Yi, et al.
Published: (2026)
by: Wei, Yi, et al.
Published: (2026)
Stochastic Difference-of-Convex Optimization with Momentum
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025)
by: Ferbach, Damien, et al.
Published: (2025)
Improved Complexity for Smooth Nonconvex Optimization: A Two-Level Online Learning Approach with Quasi-Newton Methods
by: Jiang, Ruichen, et al.
Published: (2024)
by: Jiang, Ruichen, et al.
Published: (2024)
Muon is Provably Faster with Momentum Variance Reduction
by: Qian, Xun, et al.
Published: (2025)
by: Qian, Xun, et al.
Published: (2025)
Shuffling Momentum Gradient Algorithm for Convex Optimization
by: Tran, Trang H., et al.
Published: (2024)
by: Tran, Trang H., et al.
Published: (2024)
Double Momentum Method for Lower-Level Constrained Bilevel Optimization
by: Shi, Wanli, et al.
Published: (2024)
by: Shi, Wanli, et al.
Published: (2024)
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
by: Phunyaphibarn, Prin, et al.
Published: (2023)
by: Phunyaphibarn, Prin, et al.
Published: (2023)
Similar Items
-
Efficient Sign-Based Optimization: Accelerating Convergence via Variance Reduction
by: Jiang, Wei, et al.
Published: (2024) -
Projection-Free Variance Reduction Methods for Stochastic Constrained Multi-Level Compositional Optimization
by: Jiang, Wei, et al.
Published: (2024) -
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
by: Tao, Hongyi, et al.
Published: (2026) -
Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise
by: Yu, Dingzhi, et al.
Published: (2026) -
Adaptive Variance Reduction for Stochastic Optimization under Weaker Assumptions
by: Jiang, Wei, et al.
Published: (2024)