Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Umeda, Hikaru, Iiduka, Hideaki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911095463084032
author Umeda, Hikaru
Iiduka, Hideaki
author_facet Umeda, Hikaru
Iiduka, Hideaki
contents The unprecedented growth of deep learning models has enabled remarkable advances but introduced substantial computational bottlenecks. A key factor contributing to training efficiency is batch-size and learning-rate scheduling in stochastic gradient methods. However, naive scheduling of these hyperparameters can degrade optimization efficiency and compromise generalization. Motivated by recent theoretical insights, we investigated how the batch size and learning rate should be increased during training to balance efficiency and convergence. We analyzed this problem on the basis of stochastic first-order oracle (SFO) complexity, defined as the expected number of gradient evaluations needed to reach an $ε$-approximate stationary point of the empirical loss. We theoretically derived optimal growth schedules for the batch size and learning rate that reduce SFO complexity and validated them through extensive experiments. Our results offer both theoretical insights and practical guidelines for scalable and efficient large-batch training in deep learning.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05297
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
Umeda, Hikaru
Iiduka, Hideaki
Machine Learning
Optimization and Control
The unprecedented growth of deep learning models has enabled remarkable advances but introduced substantial computational bottlenecks. A key factor contributing to training efficiency is batch-size and learning-rate scheduling in stochastic gradient methods. However, naive scheduling of these hyperparameters can degrade optimization efficiency and compromise generalization. Motivated by recent theoretical insights, we investigated how the batch size and learning rate should be increased during training to balance efficiency and convergence. We analyzed this problem on the basis of stochastic first-order oracle (SFO) complexity, defined as the expected number of gradient evaluations needed to reach an $ε$-approximate stationary point of the empirical loss. We theoretically derived optimal growth schedules for the batch size and learning rate that reduce SFO complexity and validated them through extensive experiments. Our results offer both theoretical insights and practical guidelines for scalable and efficient large-batch training in deep learning.
title Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2508.05297