AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
Fuente:
arXiv
Guardado en:
| Autores principales: | Lau, Tim Tsz-Kit, Liu, Han, Kolar, Mladen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
por: Ostroukhov, Petr, et al.
Publicado: (2024)
por: Ostroukhov, Petr, et al.
Publicado: (2024)
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2025)
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2025)
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
por: Bojovic, Matia, et al.
Publicado: (2026)
por: Bojovic, Matia, et al.
Publicado: (2026)
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
por: Liu, Zijian
Publicado: (2026)
por: Liu, Zijian
Publicado: (2026)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
por: Zhang, Minxin, et al.
Publicado: (2025)
por: Zhang, Minxin, et al.
Publicado: (2025)
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2026)
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2026)
AdaGrad under Anisotropic Smoothness
por: Liu, Yuxing, et al.
Publicado: (2024)
por: Liu, Yuxing, et al.
Publicado: (2024)
Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters
por: Yukhimchuk, Alexander, et al.
Publicado: (2026)
por: Yukhimchuk, Alexander, et al.
Publicado: (2026)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
por: Islamov, Rustem, et al.
Publicado: (2026)
por: Islamov, Rustem, et al.
Publicado: (2026)
Revisiting Convergence of AdaGrad with Relaxed Assumptions
por: Hong, Yusu, et al.
Publicado: (2024)
por: Hong, Yusu, et al.
Publicado: (2024)
GeoAdaLer: Geometric Insights into Adaptive Stochastic Gradient Descent Algorithms
por: Eleh, Chinedu, et al.
Publicado: (2024)
por: Eleh, Chinedu, et al.
Publicado: (2024)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
por: Umeda, Hikaru, et al.
Publicado: (2025)
por: Umeda, Hikaru, et al.
Publicado: (2025)
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
por: Köhne, Frederik, et al.
Publicado: (2023)
por: Köhne, Frederik, et al.
Publicado: (2023)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
por: Chezhegov, Savelii, et al.
Publicado: (2024)
por: Chezhegov, Savelii, et al.
Publicado: (2024)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
por: Oowada, Kanata, et al.
Publicado: (2025)
por: Oowada, Kanata, et al.
Publicado: (2025)
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
por: Mishkin, Aaron, et al.
Publicado: (2024)
por: Mishkin, Aaron, et al.
Publicado: (2024)
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
por: Zhou, Dongruo, et al.
Publicado: (2018)
por: Zhou, Dongruo, et al.
Publicado: (2018)
Towards Simple and Provable Parameter-Free Adaptive Gradient Methods
por: Tao, Yuanzhe, et al.
Publicado: (2024)
por: Tao, Yuanzhe, et al.
Publicado: (2024)
Modeling AdaGrad, RMSProp, and Adam with Integro-Differential Equations
por: Heredia, Carlos
Publicado: (2024)
por: Heredia, Carlos
Publicado: (2024)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
por: Umeda, Hikaru, et al.
Publicado: (2024)
por: Umeda, Hikaru, et al.
Publicado: (2024)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
por: Choudhury, Sayantan, et al.
Publicado: (2026)
por: Choudhury, Sayantan, et al.
Publicado: (2026)
Randomized Feasibility Methods for Constrained Optimization with Adaptive Step Sizes
por: Chakraborty, Abhishek, et al.
Publicado: (2026)
por: Chakraborty, Abhishek, et al.
Publicado: (2026)
AdaFisher: Adaptive Second Order Optimization via Fisher Information
por: Gomes, Damien Martins, et al.
Publicado: (2024)
por: Gomes, Damien Martins, et al.
Publicado: (2024)
Fully Stochastic Trust-Region Sequential Quadratic Programming for Equality-Constrained Optimization Problems
por: Fang, Yuchen, et al.
Publicado: (2022)
por: Fang, Yuchen, et al.
Publicado: (2022)
Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
por: Jiang, Ruichen, et al.
Publicado: (2024)
por: Jiang, Ruichen, et al.
Publicado: (2024)
TiAda: A Time-scale Adaptive Algorithm for Nonconvex Minimax Optimization
por: Li, Xiang, et al.
Publicado: (2022)
por: Li, Xiang, et al.
Publicado: (2022)
ASGO: Adaptive Structured Gradient Optimization
por: An, Kang, et al.
Publicado: (2025)
por: An, Kang, et al.
Publicado: (2025)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
por: Orvieto, Antonio, et al.
Publicado: (2024)
por: Orvieto, Antonio, et al.
Publicado: (2024)
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
por: Borodich, Ekaterina, et al.
Publicado: (2025)
por: Borodich, Ekaterina, et al.
Publicado: (2025)
AdaSwitch: An Adaptive Switching Meta-Algorithm for Learning-Augmented Bounded-Influence Problems
por: Chen, Xi, et al.
Publicado: (2025)
por: Chen, Xi, et al.
Publicado: (2025)
Adaptive Conditional Gradient Descent
por: Khademi, Abbas, et al.
Publicado: (2025)
por: Khademi, Abbas, et al.
Publicado: (2025)
GradPower: Powering Gradients for Faster Language Model Pre-Training
por: Wang, Jinbo, et al.
Publicado: (2025)
por: Wang, Jinbo, et al.
Publicado: (2025)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
por: Choudhury, Sayantan, et al.
Publicado: (2024)
por: Choudhury, Sayantan, et al.
Publicado: (2024)
Interpreting Adaptive Gradient Methods by Parameter Scaling for Learning-Rate-Free Optimization
por: Suh, Min-Kook, et al.
Publicado: (2024)
por: Suh, Min-Kook, et al.
Publicado: (2024)
Adaptive Proximal Gradient Method for Convex Optimization
por: Malitsky, Yura, et al.
Publicado: (2023)
por: Malitsky, Yura, et al.
Publicado: (2023)
Stochastic Gradient Descent with Adaptive Data
por: Che, Ethan, et al.
Publicado: (2024)
por: Che, Ethan, et al.
Publicado: (2024)
Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
por: Zhang, Dake, et al.
Publicado: (2024)
por: Zhang, Dake, et al.
Publicado: (2024)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
por: Tsukada, Yuki, et al.
Publicado: (2023)
por: Tsukada, Yuki, et al.
Publicado: (2023)
Ejemplares similares
-
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024) -
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024) -
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
por: Ostroukhov, Petr, et al.
Publicado: (2024) -
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2025) -
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
por: Bojovic, Matia, et al.
Publicado: (2026)