Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kondo, Yuichi, Iiduka, Hideaki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025)
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
Convergence Bound and Critical Batch Size of Muon Optimizer
von: Sato, Naoki, et al.
Veröffentlicht: (2025)
von: Sato, Naoki, et al.
Veröffentlicht: (2025)
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
von: Imaizumi, Kento, et al.
Veröffentlicht: (2025)
von: Imaizumi, Kento, et al.
Veröffentlicht: (2025)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
von: Tsukada, Yuki, et al.
Veröffentlicht: (2023)
von: Tsukada, Yuki, et al.
Veröffentlicht: (2023)
Iteration and Stochastic First-order Oracle Complexities of Stochastic Gradient Descent using Constant and Decaying Learning Rates
von: Imaizumi, Kento, et al.
Veröffentlicht: (2024)
von: Imaizumi, Kento, et al.
Veröffentlicht: (2024)
Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach
von: Sato, Naoki, et al.
Veröffentlicht: (2026)
von: Sato, Naoki, et al.
Veröffentlicht: (2026)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
von: Nagashima, Shuntaro, et al.
Veröffentlicht: (2026)
von: Nagashima, Shuntaro, et al.
Veröffentlicht: (2026)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
von: Iiduka, Hideaki
Veröffentlicht: (2026)
von: Iiduka, Hideaki
Veröffentlicht: (2026)
Convergence Analysis of SGD under Expected Smoothness
von: Kawamoto, Yuta, et al.
Veröffentlicht: (2025)
von: Kawamoto, Yuta, et al.
Veröffentlicht: (2025)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
von: Sato, Naoki, et al.
Veröffentlicht: (2023)
von: Sato, Naoki, et al.
Veröffentlicht: (2023)
Explicit and Implicit Graduated Optimization in Deep Neural Networks
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
von: Shen, Yikang, et al.
Veröffentlicht: (2024)
von: Shen, Yikang, et al.
Veröffentlicht: (2024)
Scaled Conjugate Gradient Method for Nonconvex Optimization in Deep Neural Networks
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
Mini-Batch Stochastic Halpern Algorithm for Nonexpansive Fixed point Problems
von: Iiduka, Hideaki
Veröffentlicht: (2026)
von: Iiduka, Hideaki
Veröffentlicht: (2026)
Smaller Batches, Bigger Gains? Investigating the Impact of Batch Sizes on Reinforcement Learning Based Real-World Production Scheduling
von: Müller, Arthur, et al.
Veröffentlicht: (2024)
von: Müller, Arthur, et al.
Veröffentlicht: (2024)
Mini-Batch Stochastic Krasnosel'ski\uı-Mann Algorithm for Nonexpansive Fixed Point Problems
von: Iiduka, Hideaki
Veröffentlicht: (2026)
von: Iiduka, Hideaki
Veröffentlicht: (2026)
Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling
von: Li, Shuaipeng, et al.
Veröffentlicht: (2024)
von: Li, Shuaipeng, et al.
Veröffentlicht: (2024)
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
von: Filatov, Oleg, et al.
Veröffentlicht: (2024)
von: Filatov, Oleg, et al.
Veröffentlicht: (2024)
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
von: Wang, Jinbo, et al.
Veröffentlicht: (2026)
von: Wang, Jinbo, et al.
Veröffentlicht: (2026)
One Size Does Not Fit All: Architecture-Aware Adaptive Batch Scheduling with DEBA
von: Belias, François, et al.
Veröffentlicht: (2025)
von: Belias, François, et al.
Veröffentlicht: (2025)
Full-Graph vs. Mini-Batch Training: Comprehensive Analysis from a Batch Size and Fan-Out Size Perspective
von: Liu, Mengfan, et al.
Veröffentlicht: (2026)
von: Liu, Mengfan, et al.
Veröffentlicht: (2026)
On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond
von: Wang, Bohan, et al.
Veröffentlicht: (2024)
von: Wang, Bohan, et al.
Veröffentlicht: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
DIVEBATCH: Accelerating Model Training Through Gradient-Diversity Aware Batch Size Adaptation
von: Chen, Yuen, et al.
Veröffentlicht: (2025)
von: Chen, Yuen, et al.
Veröffentlicht: (2025)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
von: Ostroukhov, Petr, et al.
Veröffentlicht: (2024)
von: Ostroukhov, Petr, et al.
Veröffentlicht: (2024)
Collaborative Batch Size Optimization for Federated Learning
von: Geimer, Arno, et al.
Veröffentlicht: (2025)
von: Geimer, Arno, et al.
Veröffentlicht: (2025)
Decoupled Relative Learning Rate Schedules
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
Actionable Interpretability via Causal Hypergraphs: Unravelling Batch Size Effects in Deep Learning
von: Sun, Zhongtian, et al.
Veröffentlicht: (2025)
von: Sun, Zhongtian, et al.
Veröffentlicht: (2025)
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
von: Merrill, William, et al.
Veröffentlicht: (2025)
von: Merrill, William, et al.
Veröffentlicht: (2025)
Diversified Batch Selection for Training Acceleration
von: Hong, Feng, et al.
Veröffentlicht: (2024)
von: Hong, Feng, et al.
Veröffentlicht: (2024)
The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation
von: Stewart, Lawrence, et al.
Veröffentlicht: (2024)
von: Stewart, Lawrence, et al.
Veröffentlicht: (2024)
A Concise Lyapunov Analysis of Nesterov's Accelerated Gradient Method
von: Liu, Jun
Veröffentlicht: (2025)
von: Liu, Jun
Veröffentlicht: (2025)
An Adaptive Volatility-based Learning Rate Scheduler
von: Ren, Kieran Chai Kai
Veröffentlicht: (2025)
von: Ren, Kieran Chai Kai
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025) -
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024) -
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025) -
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
von: Harada, Hinata, et al.
Veröffentlicht: (2024) -
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025)