Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
Fuente:
arXiv
Saved in:
| Main Authors: | Meterez, Alexandru, Morwani, Depen, Wu, Jingfeng, Oncescu, Costin-Andrei, Pehlevan, Cengiz, Kakade, Sham |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024)
by: Morwani, Depen, et al.
Published: (2024)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025)
by: Morwani, Depen, et al.
Published: (2025)
Feature emergence via margin maximization: case studies in algebraic tasks
by: Morwani, Depen, et al.
Published: (2023)
by: Morwani, Depen, et al.
Published: (2023)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
by: Abreu, Natalie, et al.
Published: (2025)
by: Abreu, Natalie, et al.
Published: (2025)
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
Constraint Programming Models For Serial Batch Scheduling With Minimum Batch Size
by: Huertas, Jorge A., et al.
Published: (2025)
by: Huertas, Jorge A., et al.
Published: (2025)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024)
by: Umeda, Hikaru, et al.
Published: (2024)
Convex Relaxation for Solving Large-Margin Classifiers in Hyperbolic Space
by: Yang, Sheng, et al.
Published: (2024)
by: Yang, Sheng, et al.
Published: (2024)
Deconstructing What Makes a Good Optimizer for Language Models
by: Zhao, Rosie, et al.
Published: (2024)
by: Zhao, Rosie, et al.
Published: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
by: Kou, Yiwen, et al.
Published: (2024)
by: Kou, Yiwen, et al.
Published: (2024)
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
by: Wang, Jinbo, et al.
Published: (2026)
by: Wang, Jinbo, et al.
Published: (2026)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
by: Harada, Hinata, et al.
Published: (2024)
by: Harada, Hinata, et al.
Published: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Anytime Training with Schedule-Free Spectral Optimization
by: Apte, Anuj, et al.
Published: (2026)
by: Apte, Anuj, et al.
Published: (2026)
Parallel Batch Scheduling With Incompatible Job Families Via Constraint Programming
by: Huertas, Jorge A., et al.
Published: (2024)
by: Huertas, Jorge A., et al.
Published: (2024)
Adaptive Open-Loop Step-Sizes for Accelerated Convergence Rates of the Frank-Wolfe Algorithm
by: Wirth, Elias, et al.
Published: (2025)
by: Wirth, Elias, et al.
Published: (2025)
PyCSP3-Scheduling: A Scheduling Extension for PyCSP3
by: Afifi, Sohaib
Published: (2026)
by: Afifi, Sohaib
Published: (2026)
Deep Reinforcement Learning for Flexible Job Shop Scheduling with Random Job Arrivals
by: Tang, Yu, et al.
Published: (2026)
by: Tang, Yu, et al.
Published: (2026)
On the Batch Size Selection in Stochastic Gradient Methods Using No-Replacement Sampling
by: Boresta, Marco, et al.
Published: (2025)
by: Boresta, Marco, et al.
Published: (2025)
Accelerated Affine-Invariant Convergence Rates of the Frank-Wolfe Algorithm with Open-Loop Step-Sizes
by: Wirth, Elias, et al.
Published: (2023)
by: Wirth, Elias, et al.
Published: (2023)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
Accelerating RLHF Training with Reward Variance Increase
by: Yang, Zonglin, et al.
Published: (2025)
by: Yang, Zonglin, et al.
Published: (2025)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
by: Schaipp, Fabian, et al.
Published: (2025)
by: Schaipp, Fabian, et al.
Published: (2025)
Learning-Guided Rolling Horizon Optimization for Long-Horizon Flexible Job-Shop Scheduling
by: Li, Sirui, et al.
Published: (2025)
by: Li, Sirui, et al.
Published: (2025)
MUSIC: Accelerated Convergence for Distributed Optimization With Inexact and Exact Methods
by: Wu, Mou, et al.
Published: (2024)
by: Wu, Mou, et al.
Published: (2024)
Scheduling Aerial Vehicles in an Urban Air Mobility Scheme
by: Rigas, Emmanouil S., et al.
Published: (2021)
by: Rigas, Emmanouil S., et al.
Published: (2021)
Learning Rate Schedules in the Presence of Distribution Shift
by: Fahrbach, Matthew, et al.
Published: (2023)
by: Fahrbach, Matthew, et al.
Published: (2023)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Accelerated Gradient Descent by Concatenation of Stepsize Schedules
by: Zhang, Zehao, et al.
Published: (2024)
by: Zhang, Zehao, et al.
Published: (2024)
ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule
by: Huang, Yilie, et al.
Published: (2026)
by: Huang, Yilie, et al.
Published: (2026)
Energy-Aware Model Predictive Control for Batch Manufacturing System Scheduling Under Different Electricity Pricing Strategies
by: Li, Hongliang, et al.
Published: (2025)
by: Li, Hongliang, et al.
Published: (2025)
Similar Items
-
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2025) -
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026) -
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024) -
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
by: Oncescu, Costin-Andrei, et al.
Published: (2026) -
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024)