AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
Fuente:
arXiv
Saved in:
| Main Authors: | Ostroukhov, Petr, Zhumabayeva, Aigerim, Xiang, Chulu, Gasnikov, Alexander, Takáč, Martin, Kamzolov, Dmitry |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
by: Abdukhakimov, Farshed, et al.
Published: (2023)
by: Abdukhakimov, Farshed, et al.
Published: (2023)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Exploring Jacobian Inexactness in Second-Order Methods for Variational Inequalities: Lower Bounds, Optimal Algorithms and Quasi-Newton Approximations
by: Agafonov, Artem, et al.
Published: (2024)
by: Agafonov, Artem, et al.
Published: (2024)
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Methods with Local Steps and Random Reshuffling for Generally Smooth Non-Convex Federated Optimization
by: Demidovich, Yury, et al.
Published: (2024)
by: Demidovich, Yury, et al.
Published: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
OPTAMI: Global Superlinear Convergence of High-order Methods
by: Kamzolov, Dmitry, et al.
Published: (2024)
by: Kamzolov, Dmitry, et al.
Published: (2024)
Adaptive Regularized Newton Method with Inexact Hessian
by: Shestakov, Aleksandr, et al.
Published: (2025)
by: Shestakov, Aleksandr, et al.
Published: (2025)
Exploiting higher-order derivatives in convex optimization methods
by: Kamzolov, Dmitry, et al.
Published: (2022)
by: Kamzolov, Dmitry, et al.
Published: (2022)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
by: Bojovic, Matia, et al.
Published: (2026)
by: Bojovic, Matia, et al.
Published: (2026)
$ψ$DAG: Projected Stochastic Approximation Iteration for DAG Structure Learning
by: Ziu, Klea, et al.
Published: (2024)
by: Ziu, Klea, et al.
Published: (2024)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
by: Oowada, Kanata, et al.
Published: (2025)
by: Oowada, Kanata, et al.
Published: (2025)
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
by: Köhne, Frederik, et al.
Published: (2023)
by: Köhne, Frederik, et al.
Published: (2023)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
by: Choudhury, Sayantan, et al.
Published: (2024)
by: Choudhury, Sayantan, et al.
Published: (2024)
Stochastic Gradient Methods with Preconditioned Updates
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024)
by: Umeda, Hikaru, et al.
Published: (2024)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
by: Tsukada, Yuki, et al.
Published: (2023)
by: Tsukada, Yuki, et al.
Published: (2023)
Randomized Feasibility Methods for Constrained Optimization with Adaptive Step Sizes
by: Chakraborty, Abhishek, et al.
Published: (2026)
by: Chakraborty, Abhishek, et al.
Published: (2026)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Similarity, Compression and Local Steps: Three Pillars of Efficient Communications for Distributed Variational Inequalities
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
by: Harada, Hinata, et al.
Published: (2024)
by: Harada, Hinata, et al.
Published: (2024)
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
by: Imaizumi, Kento, et al.
Published: (2025)
by: Imaizumi, Kento, et al.
Published: (2025)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024)
by: Liu, Yuxing, et al.
Published: (2024)
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
by: Wang, Jinbo, et al.
Published: (2026)
by: Wang, Jinbo, et al.
Published: (2026)
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
by: Liu, Zijian
Published: (2026)
by: Liu, Zijian
Published: (2026)
Revisiting Convergence of AdaGrad with Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024)
by: Hong, Yusu, et al.
Published: (2024)
Advancing the lower bounds: An accelerated, stochastic, second-order method with optimal adaptation to inexactness
by: Agafonov, Artem, et al.
Published: (2023)
by: Agafonov, Artem, et al.
Published: (2023)
Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes
by: Metel, Michael R.
Published: (2022)
by: Metel, Michael R.
Published: (2022)
Simple Stepsize for Quasi-Newton Methods with Global Convergence Guarantees
by: Agafonov, Artem, et al.
Published: (2025)
by: Agafonov, Artem, et al.
Published: (2025)
Decentralized Inexact Cubic Newton Method with Consensus Procedure
by: Agafonov, Artem, et al.
Published: (2026)
by: Agafonov, Artem, et al.
Published: (2026)
TiAda: A Time-scale Adaptive Algorithm for Nonconvex Minimax Optimization
by: Li, Xiang, et al.
Published: (2022)
by: Li, Xiang, et al.
Published: (2022)
Online Linear Programming with Batching
by: Xu, Haoran, et al.
Published: (2024)
by: Xu, Haoran, et al.
Published: (2024)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026)
by: Riabinin, Artem, et al.
Published: (2026)
Lower Bounds and Optimal Algorithms for Non-Smooth Convex Decentralized Optimization over Time-Varying Networks
by: Kovalev, Dmitry, et al.
Published: (2024)
by: Kovalev, Dmitry, et al.
Published: (2024)
Similar Items
-
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024) -
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
by: Abdukhakimov, Farshed, et al.
Published: (2023) -
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024) -
Exploring Jacobian Inexactness in Second-Order Methods for Variational Inequalities: Lower Bounds, Optimal Algorithms and Quasi-Newton Approximations
by: Agafonov, Artem, et al.
Published: (2024) -
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)