Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
Fuente:
arXiv
Saved in:
| Main Authors: | Umeda, Hikaru, Iiduka, Hideaki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024)
by: Umeda, Hikaru, et al.
Published: (2024)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
by: Harada, Hinata, et al.
Published: (2024)
by: Harada, Hinata, et al.
Published: (2024)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
by: Oowada, Kanata, et al.
Published: (2025)
by: Oowada, Kanata, et al.
Published: (2025)
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
by: Imaizumi, Kento, et al.
Published: (2025)
by: Imaizumi, Kento, et al.
Published: (2025)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
by: Tsukada, Yuki, et al.
Published: (2023)
by: Tsukada, Yuki, et al.
Published: (2023)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
by: Sato, Naoki, et al.
Published: (2024)
by: Sato, Naoki, et al.
Published: (2024)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
by: Iiduka, Hideaki
Published: (2026)
by: Iiduka, Hideaki
Published: (2026)
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
by: Sato, Naoki, et al.
Published: (2023)
by: Sato, Naoki, et al.
Published: (2023)
Mini-Batch Stochastic Halpern Algorithm for Nonexpansive Fixed point Problems
by: Iiduka, Hideaki
Published: (2026)
by: Iiduka, Hideaki
Published: (2026)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
by: Kondo, Yuichi, et al.
Published: (2025)
by: Kondo, Yuichi, et al.
Published: (2025)
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
by: Wang, Jinbo, et al.
Published: (2026)
by: Wang, Jinbo, et al.
Published: (2026)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
by: Tyurin, Alexander, et al.
Published: (2024)
by: Tyurin, Alexander, et al.
Published: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)
by: Wang, Runzhe, et al.
Published: (2023)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Mini-Batch Stochastic Krasnosel'ski\uı-Mann Algorithm for Nonexpansive Fixed Point Problems
by: Iiduka, Hideaki
Published: (2026)
by: Iiduka, Hideaki
Published: (2026)
SGD for Variational Inference: Tackling Unbounded Variance via Preconditioning and Dynamic Batching
by: Labarrière, Hippolyte, et al.
Published: (2026)
by: Labarrière, Hippolyte, et al.
Published: (2026)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization
by: Zhang, Haihan, et al.
Published: (2024)
by: Zhang, Haihan, et al.
Published: (2024)
Optimal Projection-Free Adaptive SGD for Matrix Optimization
by: Kovalev, Dmitry
Published: (2026)
by: Kovalev, Dmitry
Published: (2026)
AutoSGD: Automatic Learning Rate Selection for Stochastic Gradient Descent
by: Surjanovic, Nikola, et al.
Published: (2025)
by: Surjanovic, Nikola, et al.
Published: (2025)
Learning Rate Schedules in the Presence of Distribution Shift
by: Fahrbach, Matthew, et al.
Published: (2023)
by: Fahrbach, Matthew, et al.
Published: (2023)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
by: Kovalev, Dmitry, et al.
Published: (2025)
by: Kovalev, Dmitry, et al.
Published: (2025)
A Minibatch-SGD-Based Learning Meta-Policy for Inventory Systems with Myopic Optimal Policy
by: Lyu, Jiameng, et al.
Published: (2024)
by: Lyu, Jiameng, et al.
Published: (2024)
General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
by: Wu, Haotian
Published: (2025)
by: Wu, Haotian
Published: (2025)
Bias-Optimal Bounds for SGD: A Computer-Aided Lyapunov Analysis
by: Cortild, Daniel, et al.
Published: (2025)
by: Cortild, Daniel, et al.
Published: (2025)
Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes
by: Metel, Michael R.
Published: (2022)
by: Metel, Michael R.
Published: (2022)
ROOT-SGD: Sharp Nonasymptotics and Near-Optimal Asymptotics in a Single Algorithm
by: Li, Chris Junchi, et al.
Published: (2020)
by: Li, Chris Junchi, et al.
Published: (2020)
Sharp High-Probability Rates for Nonlinear SGD under Heavy-Tailed Noise via Symmetrization
by: Armacki, Aleksandar, et al.
Published: (2025)
by: Armacki, Aleksandar, et al.
Published: (2025)
Ringmaster ASGD: The First Asynchronous SGD with Optimal Time Complexity
by: Maranjyan, Artavazd, et al.
Published: (2025)
by: Maranjyan, Artavazd, et al.
Published: (2025)
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Stochastic Variance-Reduced Newton: Accelerating Finite-Sum Minimization with Large Batches
by: Dereziński, Michał
Published: (2022)
by: Dereziński, Michał
Published: (2022)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
by: Kamo, Keisuke, et al.
Published: (2025)
by: Kamo, Keisuke, et al.
Published: (2025)
Similar Items
-
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
by: Umeda, Hikaru, et al.
Published: (2025) -
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024) -
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
by: Harada, Hinata, et al.
Published: (2024) -
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
by: Oowada, Kanata, et al.
Published: (2025) -
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
by: Imaizumi, Kento, et al.
Published: (2025)