AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Lau, Tim Tsz-Kit, Liu, Han, Kolar, Mladen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
by: Bojovic, Matia, et al.
Published: (2026)
by: Bojovic, Matia, et al.
Published: (2026)
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
by: Liu, Zijian
Published: (2026)
by: Liu, Zijian
Published: (2026)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
by: Lau, Tim Tsz-Kit, et al.
Published: (2026)
by: Lau, Tim Tsz-Kit, et al.
Published: (2026)
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024)
by: Liu, Yuxing, et al.
Published: (2024)
Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters
by: Yukhimchuk, Alexander, et al.
Published: (2026)
by: Yukhimchuk, Alexander, et al.
Published: (2026)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Revisiting Convergence of AdaGrad with Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024)
by: Hong, Yusu, et al.
Published: (2024)
GeoAdaLer: Geometric Insights into Adaptive Stochastic Gradient Descent Algorithms
by: Eleh, Chinedu, et al.
Published: (2024)
by: Eleh, Chinedu, et al.
Published: (2024)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
by: Köhne, Frederik, et al.
Published: (2023)
by: Köhne, Frederik, et al.
Published: (2023)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
by: Oowada, Kanata, et al.
Published: (2025)
by: Oowada, Kanata, et al.
Published: (2025)
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
by: Zhou, Dongruo, et al.
Published: (2018)
by: Zhou, Dongruo, et al.
Published: (2018)
Towards Simple and Provable Parameter-Free Adaptive Gradient Methods
by: Tao, Yuanzhe, et al.
Published: (2024)
by: Tao, Yuanzhe, et al.
Published: (2024)
Modeling AdaGrad, RMSProp, and Adam with Integro-Differential Equations
by: Heredia, Carlos
Published: (2024)
by: Heredia, Carlos
Published: (2024)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024)
by: Umeda, Hikaru, et al.
Published: (2024)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
by: Choudhury, Sayantan, et al.
Published: (2026)
by: Choudhury, Sayantan, et al.
Published: (2026)
Randomized Feasibility Methods for Constrained Optimization with Adaptive Step Sizes
by: Chakraborty, Abhishek, et al.
Published: (2026)
by: Chakraborty, Abhishek, et al.
Published: (2026)
AdaFisher: Adaptive Second Order Optimization via Fisher Information
by: Gomes, Damien Martins, et al.
Published: (2024)
by: Gomes, Damien Martins, et al.
Published: (2024)
Fully Stochastic Trust-Region Sequential Quadratic Programming for Equality-Constrained Optimization Problems
by: Fang, Yuchen, et al.
Published: (2022)
by: Fang, Yuchen, et al.
Published: (2022)
Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
by: Jiang, Ruichen, et al.
Published: (2024)
by: Jiang, Ruichen, et al.
Published: (2024)
TiAda: A Time-scale Adaptive Algorithm for Nonconvex Minimax Optimization
by: Li, Xiang, et al.
Published: (2022)
by: Li, Xiang, et al.
Published: (2022)
ASGO: Adaptive Structured Gradient Optimization
by: An, Kang, et al.
Published: (2025)
by: An, Kang, et al.
Published: (2025)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
by: Orvieto, Antonio, et al.
Published: (2024)
by: Orvieto, Antonio, et al.
Published: (2024)
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
by: Borodich, Ekaterina, et al.
Published: (2025)
by: Borodich, Ekaterina, et al.
Published: (2025)
AdaSwitch: An Adaptive Switching Meta-Algorithm for Learning-Augmented Bounded-Influence Problems
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Adaptive Conditional Gradient Descent
by: Khademi, Abbas, et al.
Published: (2025)
by: Khademi, Abbas, et al.
Published: (2025)
GradPower: Powering Gradients for Faster Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
by: Choudhury, Sayantan, et al.
Published: (2024)
by: Choudhury, Sayantan, et al.
Published: (2024)
Interpreting Adaptive Gradient Methods by Parameter Scaling for Learning-Rate-Free Optimization
by: Suh, Min-Kook, et al.
Published: (2024)
by: Suh, Min-Kook, et al.
Published: (2024)
Adaptive Proximal Gradient Method for Convex Optimization
by: Malitsky, Yura, et al.
Published: (2023)
by: Malitsky, Yura, et al.
Published: (2023)
Stochastic Gradient Descent with Adaptive Data
by: Che, Ethan, et al.
Published: (2024)
by: Che, Ethan, et al.
Published: (2024)
Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
by: Zhang, Dake, et al.
Published: (2024)
by: Zhang, Dake, et al.
Published: (2024)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
by: Tsukada, Yuki, et al.
Published: (2023)
by: Tsukada, Yuki, et al.
Published: (2023)
Similar Items
-
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024) -
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024) -
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024) -
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
by: Lau, Tim Tsz-Kit, et al.
Published: (2025) -
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
by: Bojovic, Matia, et al.
Published: (2026)