PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Lau, Tim Tsz-Kit, Long, Qi, Su, Weijie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
by: Lau, Tim Tsz-Kit, et al.
Published: (2026)
by: Lau, Tim Tsz-Kit, et al.
Published: (2026)
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix Factorization
by: Zhang, Gavin, et al.
Published: (2025)
by: Zhang, Gavin, et al.
Published: (2025)
Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
by: Su, Weijie
Published: (2025)
by: Su, Weijie
Published: (2025)
Efficient Over-parameterized Matrix Sensing from Noisy Measurements via Alternating Preconditioned Gradient Descent
by: Liu, Zhiyu, et al.
Published: (2025)
by: Liu, Zhiyu, et al.
Published: (2025)
Stochastic Gradient Methods with Preconditioned Updates
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
AGD: an Auto-switchable Optimizer using Stepwise Gradient Difference for Preconditioning Matrix
by: Yue, Yun, et al.
Published: (2023)
by: Yue, Yun, et al.
Published: (2023)
Mirror and Preconditioned Gradient Descent in Wasserstein Space
by: Bonet, Clément, et al.
Published: (2024)
by: Bonet, Clément, et al.
Published: (2024)
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
by: Bojovic, Matia, et al.
Published: (2026)
by: Bojovic, Matia, et al.
Published: (2026)
HiGrad: Uncertainty Quantification for Online Learning and Stochastic Approximation
by: Su, Weijie J., et al.
Published: (2018)
by: Su, Weijie J., et al.
Published: (2018)
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
by: Köhne, Frederik, et al.
Published: (2023)
by: Köhne, Frederik, et al.
Published: (2023)
GradPower: Powering Gradients for Faster Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025)
by: Kovalev, Dmitry
Published: (2025)
Guarantees of a Preconditioned Subgradient Algorithm for Overparameterized Asymmetric Low-rank Matrix Recovery
by: Giampouras, Paris, et al.
Published: (2024)
by: Giampouras, Paris, et al.
Published: (2024)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
by: Liu, Zijian
Published: (2026)
by: Liu, Zijian
Published: (2026)
Preconditioned Gradient Descent for Overparameterized Nonconvex Burer--Monteiro Factorization with Global Optimality Certification
by: Zhang, Gavin, et al.
Published: (2022)
by: Zhang, Gavin, et al.
Published: (2022)
Efficient Low-Tubal-Rank Tensor Estimation via Alternating Preconditioned Gradient Descent
by: Liu, Zhiyu, et al.
Published: (2025)
by: Liu, Zhiyu, et al.
Published: (2025)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
The Power of Preconditioning in Overparameterized Low-Rank Matrix Sensing
by: Xu, Xingyu, et al.
Published: (2023)
by: Xu, Xingyu, et al.
Published: (2023)
PROMISE: Preconditioned Stochastic Optimization Methods by Incorporating Scalable Curvature Estimates
by: Frangella, Zachary, et al.
Published: (2023)
by: Frangella, Zachary, et al.
Published: (2023)
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024)
by: Liu, Yuxing, et al.
Published: (2024)
Langevin Dynamics: A Unified Perspective on Optimization via Lyapunov Potentials
by: Chen, August Y., et al.
Published: (2024)
by: Chen, August Y., et al.
Published: (2024)
Fast and Accurate Estimation of Low-Rank Matrices from Noisy Measurements via Preconditioned Non-Convex Gradient Descent
by: Zhang, Gavin, et al.
Published: (2023)
by: Zhang, Gavin, et al.
Published: (2023)
Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided Preconditioning
by: Li, Huan, et al.
Published: (2026)
by: Li, Huan, et al.
Published: (2026)
Revisiting Convergence of AdaGrad with Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024)
by: Hong, Yusu, et al.
Published: (2024)
A Homogenization Approach for Gradient-Dominated Stochastic Optimization
by: Tan, Jiyuan, et al.
Published: (2023)
by: Tan, Jiyuan, et al.
Published: (2023)
Convergence of Alternating Gradient Descent for Matrix Factorization
by: Ward, Rachel, et al.
Published: (2023)
by: Ward, Rachel, et al.
Published: (2023)
Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
by: Jiang, Ruichen, et al.
Published: (2024)
by: Jiang, Ruichen, et al.
Published: (2024)
Tree-Preconditioned Differentiable Optimization and Axioms as Layers
by: Liao, Yuexin
Published: (2025)
by: Liao, Yuexin
Published: (2025)
Quadratic Gradient: A Unified Framework Bridging Gradient Descent and Newton-Type Methods by Synthesizing Hessians and Gradients
by: Chiang, John
Published: (2022)
by: Chiang, John
Published: (2022)
Robust Learning Rate Selection for Stochastic Optimization via Splitting Diagnostic
by: Sordello, Matteo, et al.
Published: (2019)
by: Sordello, Matteo, et al.
Published: (2019)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
A Unified Framework for Gradient Aggregation in Multi-Objective Optimization
by: Hu, Zeou, et al.
Published: (2026)
by: Hu, Zeou, et al.
Published: (2026)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Convergence of Gradient Descent with Small Initialization for Unregularized Matrix Completion
by: Ma, Jianhao, et al.
Published: (2024)
by: Ma, Jianhao, et al.
Published: (2024)
Projective Proximal Gradient Descent for A Class of Nonconvex Nonsmooth Optimization Problems: Fast Convergence Without Kurdyka-Lojasiewicz (KL) Property
by: Yang, Yingzhen, et al.
Published: (2023)
by: Yang, Yingzhen, et al.
Published: (2023)
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
by: Oikonomidis, Konstantinos, et al.
Published: (2026)
by: Oikonomidis, Konstantinos, et al.
Published: (2026)
Similar Items
-
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024) -
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
by: Lau, Tim Tsz-Kit, et al.
Published: (2026) -
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024) -
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024) -
Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix Factorization
by: Zhang, Gavin, et al.
Published: (2025)