Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jinbo, Li, Binghui, Zhou, Zhanpeng, Wang, Mingze, Sun, Yuxuan, Zhang, Jiaqi, Cai, Xunliang, Wu, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GradPower: Powering Gradients for Faster Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Muon in Associative Memory Learning: Training Dynamics and Scaling Laws
von: Li, Binghui, et al.
Veröffentlicht: (2026)
von: Li, Binghui, et al.
Veröffentlicht: (2026)
Achieving Margin Maximization Exponentially Fast via Progressive Norm Rescaling
von: Wang, Mingze, et al.
Veröffentlicht: (2023)
von: Wang, Mingze, et al.
Veröffentlicht: (2023)
Constraint Programming Models For Serial Batch Scheduling With Minimum Batch Size
von: Huertas, Jorge A., et al.
Veröffentlicht: (2025)
von: Huertas, Jorge A., et al.
Veröffentlicht: (2025)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
von: Ostroukhov, Petr, et al.
Veröffentlicht: (2024)
von: Ostroukhov, Petr, et al.
Veröffentlicht: (2024)
How Does Critical Batch Size Scale in Pre-training?
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
Improving Generalization and Convergence by Enhancing Implicit Regularization
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
Convergence and Stability of a Catching-Up Algorithm for Differential Inclusions with Maximal Monotone Operators
von: Cao, Tan H., et al.
Veröffentlicht: (2026)
von: Cao, Tan H., et al.
Veröffentlicht: (2026)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Optimal Learning-Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay
von: Li, Binghui, et al.
Veröffentlicht: (2026)
von: Li, Binghui, et al.
Veröffentlicht: (2026)
A Random Batch Method for the Efficient Simulation and Optimal Control of Networked 1-D Wave Equations
von: Veldman, Daniel, et al.
Veröffentlicht: (2025)
von: Veldman, Daniel, et al.
Veröffentlicht: (2025)
On the Batch Size Selection in Stochastic Gradient Methods Using No-Replacement Sampling
von: Boresta, Marco, et al.
Veröffentlicht: (2025)
von: Boresta, Marco, et al.
Veröffentlicht: (2025)
Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate Schedules
von: Li, Binghui, et al.
Veröffentlicht: (2025)
von: Li, Binghui, et al.
Veröffentlicht: (2025)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
Parallel Batch Scheduling With Incompatible Job Families Via Constraint Programming
von: Huertas, Jorge A., et al.
Veröffentlicht: (2024)
von: Huertas, Jorge A., et al.
Veröffentlicht: (2024)
How Transformers Get Rich: Approximation and Dynamics Analysis
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
von: Ziyin, Liu, et al.
Veröffentlicht: (2024)
von: Ziyin, Liu, et al.
Veröffentlicht: (2024)
A Network Flow Approach to Optimal Scheduling in Supply Chain Logistics
von: Wang, Yichen, et al.
Veröffentlicht: (2024)
von: Wang, Yichen, et al.
Veröffentlicht: (2024)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
A Novel State-Centric Necessary Condition for Time-Optimal Control of Controllable Linear Systems Based on Augmented Switching Laws (Extended Version)
von: Wang, Yunan, et al.
Veröffentlicht: (2024)
von: Wang, Yunan, et al.
Veröffentlicht: (2024)
Exact Quadratic Penalty Function for Symplectic Eigenvalue Problem
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
Quantum Hamiltonian Descent based Augmented Lagrangian Method for Constrained Nonconvex Nonlinear Optimization
von: Li, Mingze, et al.
Veröffentlicht: (2025)
von: Li, Mingze, et al.
Veröffentlicht: (2025)
Optimal Analysis of Method with Batching for Monotone Stochastic Finite-Sum Variational Inequalities
von: Pichugin, Alexander, et al.
Veröffentlicht: (2024)
von: Pichugin, Alexander, et al.
Veröffentlicht: (2024)
Neural Event-Triggered Control with Optimal Scheduling
von: Yang, Luan, et al.
Veröffentlicht: (2025)
von: Yang, Luan, et al.
Veröffentlicht: (2025)
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
Tight Big-Ms for Optimal Transmission Switching
von: Pineda, Salvador, et al.
Veröffentlicht: (2023)
von: Pineda, Salvador, et al.
Veröffentlicht: (2023)
Joint Order Selection, Allocation, Batching and Picking for Large Scale Warehouses
von: Abelli, Giorgio, et al.
Veröffentlicht: (2024)
von: Abelli, Giorgio, et al.
Veröffentlicht: (2024)
Batch-based Bayesian Optimal Experimental Design in Linear Inverse Problems
von: Mäkinen, Sofia, et al.
Veröffentlicht: (2026)
von: Mäkinen, Sofia, et al.
Veröffentlicht: (2026)
Turnpike Property of Mean-Field Linear-Quadratic Optimal Control Problems in Infinite-Horizon with Regime Switching
von: Mei, Hongwei, et al.
Veröffentlicht: (2025)
von: Mei, Hongwei, et al.
Veröffentlicht: (2025)
Optimal Control of Switched Systems Governed by Logical Switching Dynamics
von: Zhang, Xiao, et al.
Veröffentlicht: (2026)
von: Zhang, Xiao, et al.
Veröffentlicht: (2026)
A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation
von: Zhao, Heyang, et al.
Veröffentlicht: (2023)
von: Zhao, Heyang, et al.
Veröffentlicht: (2023)
Optimal BESS Scheduling for Multi-Market Participation in the Nordics
von: Hameed, Zeenat, et al.
Veröffentlicht: (2025)
von: Hameed, Zeenat, et al.
Veröffentlicht: (2025)
Towards Optimal Control and Algorithmic Structure of Decompression Schedules
von: Marsh, Benjamin
Veröffentlicht: (2025)
von: Marsh, Benjamin
Veröffentlicht: (2025)
Infinite Horizon Mean-Field Linear-Quadratic Optimal Control Problems with Switching and Indefinite-Weighted Costs
von: Mei, Hongwei, et al.
Veröffentlicht: (2025)
von: Mei, Hongwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GradPower: Powering Gradients for Faster Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025) -
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025) -
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025) -
Muon in Associative Memory Learning: Training Dynamics and Scaling Laws
von: Li, Binghui, et al.
Veröffentlicht: (2026) -
Achieving Margin Maximization Exponentially Fast via Progressive Norm Rescaling
von: Wang, Mingze, et al.
Veröffentlicht: (2023)