AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Minxin, Liu, Yuxuan, Schaeffer, Hayden |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
por: Zhang, Minxin, et al.
Publicado: (2026)
por: Zhang, Minxin, et al.
Publicado: (2026)
AdaGrad under Anisotropic Smoothness
por: Liu, Yuxing, et al.
Publicado: (2024)
por: Liu, Yuxing, et al.
Publicado: (2024)
Revisiting Convergence of AdaGrad with Relaxed Assumptions
por: Hong, Yusu, et al.
Publicado: (2024)
por: Hong, Yusu, et al.
Publicado: (2024)
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
por: Bojovic, Matia, et al.
Publicado: (2026)
por: Bojovic, Matia, et al.
Publicado: (2026)
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
por: Liu, Zijian
Publicado: (2026)
por: Liu, Zijian
Publicado: (2026)
Modeling AdaGrad, RMSProp, and Adam with Integro-Differential Equations
por: Heredia, Carlos
Publicado: (2024)
por: Heredia, Carlos
Publicado: (2024)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
por: Chezhegov, Savelii, et al.
Publicado: (2024)
por: Chezhegov, Savelii, et al.
Publicado: (2024)
Universality of AdaGrad Stepsizes for Stochastic Optimization: Inexact Oracle, Acceleration and Variance Reduction
por: Rodomanov, Anton, et al.
Publicado: (2024)
por: Rodomanov, Anton, et al.
Publicado: (2024)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
por: Choudhury, Sayantan, et al.
Publicado: (2024)
por: Choudhury, Sayantan, et al.
Publicado: (2024)
Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
por: Jiang, Ruichen, et al.
Publicado: (2024)
por: Jiang, Ruichen, et al.
Publicado: (2024)
A Riemannian AdaGrad-Norm Method
por: Bento, Glaydston de C., et al.
Publicado: (2025)
por: Bento, Glaydston de C., et al.
Publicado: (2025)
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
por: Ostroukhov, Petr, et al.
Publicado: (2024)
por: Ostroukhov, Petr, et al.
Publicado: (2024)
Last Iterate Convergence of AdaGrad-Norm for Convex Non-Smooth Optimization
por: Preobrazhenskaia, Margarita, et al.
Publicado: (2026)
por: Preobrazhenskaia, Margarita, et al.
Publicado: (2026)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
por: Khaled, Ahmed, et al.
Publicado: (2025)
por: Khaled, Ahmed, et al.
Publicado: (2025)
Drop-Muon: Update Less, Converge Faster
por: Gruntkowska, Kaja, et al.
Publicado: (2025)
por: Gruntkowska, Kaja, et al.
Publicado: (2025)
Beyond the Ideal: Analyzing the Inexact Muon Update
por: Shulgin, Egor, et al.
Publicado: (2025)
por: Shulgin, Egor, et al.
Publicado: (2025)
Stability and convergence analysis of AdaGrad for non-convex optimization via novel stopping time-based techniques
por: Jin, Ruinan, et al.
Publicado: (2024)
por: Jin, Ruinan, et al.
Publicado: (2024)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
por: Orvieto, Antonio, et al.
Publicado: (2024)
por: Orvieto, Antonio, et al.
Publicado: (2024)
Preconditioning Benefits of Spectral Orthogonalization in Muon
por: Ma, Jianhao, et al.
Publicado: (2026)
por: Ma, Jianhao, et al.
Publicado: (2026)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
por: Wu, Haotian
Publicado: (2025)
por: Wu, Haotian
Publicado: (2025)
MARINA-P: Superior Performance in Non-smooth Federated Optimization with Adaptive Stepsizes
por: Sokolov, Igor, et al.
Publicado: (2024)
por: Sokolov, Igor, et al.
Publicado: (2024)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
por: Zhang, Yixuan, et al.
Publicado: (2024)
por: Zhang, Yixuan, et al.
Publicado: (2024)
Stochastic Approximation with Block Coordinate Optimal Stepsizes
por: Jiang, Tao, et al.
Publicado: (2025)
por: Jiang, Tao, et al.
Publicado: (2025)
AdaFisher: Adaptive Second Order Optimization via Fisher Information
por: Gomes, Damien Martins, et al.
Publicado: (2024)
por: Gomes, Damien Martins, et al.
Publicado: (2024)
New Perspectives on the Polyak Stepsize: Surrogate Functions and Negative Results
por: Orabona, Francesco, et al.
Publicado: (2025)
por: Orabona, Francesco, et al.
Publicado: (2025)
Simple Stepsize for Quasi-Newton Methods with Global Convergence Guarantees
por: Agafonov, Artem, et al.
Publicado: (2025)
por: Agafonov, Artem, et al.
Publicado: (2025)
Bias and Extrapolation in Markovian Linear Stochastic Approximation with Constant Stepsizes
por: Huo, Dongyan, et al.
Publicado: (2022)
por: Huo, Dongyan, et al.
Publicado: (2022)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
GeoAdaLer: Geometric Insights into Adaptive Stochastic Gradient Descent Algorithms
por: Eleh, Chinedu, et al.
Publicado: (2024)
por: Eleh, Chinedu, et al.
Publicado: (2024)
TiAda: A Time-scale Adaptive Algorithm for Nonconvex Minimax Optimization
por: Li, Xiang, et al.
Publicado: (2022)
por: Li, Xiang, et al.
Publicado: (2022)
Phases of Muon: When Muon Eclipses SignSGD
por: Paquette, Elliot, et al.
Publicado: (2026)
por: Paquette, Elliot, et al.
Publicado: (2026)
LiMuon: Light and Fast Muon Optimizer for Large Models
por: Huang, Feihu, et al.
Publicado: (2025)
por: Huang, Feihu, et al.
Publicado: (2025)
Enhancing Stochastic Optimization for Statistical Efficiency Using ROOT-SGD with Diminishing Stepsize
por: Li, Chris Junchi
Publicado: (2024)
por: Li, Chris Junchi
Publicado: (2024)
AdaSwitch: An Adaptive Switching Meta-Algorithm for Learning-Augmented Bounded-Influence Problems
por: Chen, Xi, et al.
Publicado: (2025)
por: Chen, Xi, et al.
Publicado: (2025)
The Collusion of Memory and Nonlinearity in Stochastic Approximation With Constant Stepsize
por: Huo, Dongyan, et al.
Publicado: (2024)
por: Huo, Dongyan, et al.
Publicado: (2024)
Muon Optimizes Under Spectral Norm Constraints
por: Chen, Lizhang, et al.
Publicado: (2025)
por: Chen, Lizhang, et al.
Publicado: (2025)
Two-Timescale Linear Stochastic Approximation: Constant Stepsizes Go a Long Way
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
GradPower: Powering Gradients for Faster Language Model Pre-Training
por: Wang, Jinbo, et al.
Publicado: (2025)
por: Wang, Jinbo, et al.
Publicado: (2025)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
por: Roulet, Vincent, et al.
Publicado: (2023)
por: Roulet, Vincent, et al.
Publicado: (2023)
Ejemplares similares
-
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
por: Zhang, Minxin, et al.
Publicado: (2026) -
AdaGrad under Anisotropic Smoothness
por: Liu, Yuxing, et al.
Publicado: (2024) -
Revisiting Convergence of AdaGrad with Relaxed Assumptions
por: Hong, Yusu, et al.
Publicado: (2024) -
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
por: Bojovic, Matia, et al.
Publicado: (2026) -
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
por: Liu, Zijian
Publicado: (2026)