Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meterez, Alexandru, Morwani, Depen, Wu, Jingfeng, Oncescu, Costin-Andrei, Pehlevan, Cengiz, Kakade, Sham
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912652787187712
author Meterez, Alexandru
Morwani, Depen
Wu, Jingfeng
Oncescu, Costin-Andrei
Pehlevan, Cengiz
Kakade, Sham
author_facet Meterez, Alexandru
Morwani, Depen
Wu, Jingfeng
Oncescu, Costin-Andrei
Pehlevan, Cengiz
Kakade, Sham
contents Increasing the batch size during training -- a ''batch ramp'' -- is a promising strategy to accelerate large language model pretraining. While for SGD, doubling the batch size can be equivalent to halving the learning rate, the optimal strategy for adaptive optimizers like Adam is less clear. As a result, any batch-ramp scheduling, if used at all, is typically tuned heuristically. This work develops a principled framework for batch-size scheduling and introduces Seesaw: whenever a standard scheduler would halve the learning rate, Seesaw instead multiplies it by $1/\sqrt{2}$ and doubles the batch size, preserving loss dynamics while reducing serial steps. Theoretically, we provide, to our knowledge, the first finite-sample proof of equivalence between learning-rate decay and batch-size ramp-up for SGD on noisy linear regression, and we extend this equivalence to normalized SGD, a tractable proxy for Adam, under a variance-dominated regime observed in practice. Empirically, on 150M/300M/600M-parameter models trained at Chinchilla scale using a constant (critical) batch size, Seesaw matches cosine decay at equal FLOPs while reducing wall-clock time by $\approx 36\%$, approaching the theoretical limit implied by our analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14717
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
Meterez, Alexandru
Morwani, Depen
Wu, Jingfeng
Oncescu, Costin-Andrei
Pehlevan, Cengiz
Kakade, Sham
Machine Learning
Artificial Intelligence
Optimization and Control
Increasing the batch size during training -- a ''batch ramp'' -- is a promising strategy to accelerate large language model pretraining. While for SGD, doubling the batch size can be equivalent to halving the learning rate, the optimal strategy for adaptive optimizers like Adam is less clear. As a result, any batch-ramp scheduling, if used at all, is typically tuned heuristically. This work develops a principled framework for batch-size scheduling and introduces Seesaw: whenever a standard scheduler would halve the learning rate, Seesaw instead multiplies it by $1/\sqrt{2}$ and doubles the batch size, preserving loss dynamics while reducing serial steps. Theoretically, we provide, to our knowledge, the first finite-sample proof of equivalence between learning-rate decay and batch-size ramp-up for SGD on noisy linear regression, and we extend this equivalence to normalized SGD, a tractable proxy for Adam, under a variance-dominated regime observed in practice. Empirically, on 150M/300M/600M-parameter models trained at Chinchilla scale using a constant (critical) batch size, Seesaw matches cosine decay at equal FLOPs while reducing wall-clock time by $\approx 36\%$, approaching the theoretical limit implied by our analysis.
title Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2510.14717