How Does Critical Batch Size Scale in Pre-training?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Hanlin, Morwani, Depen, Vyas, Nikhil, Wu, Jingfeng, Zou, Difan, Ghai, Udaya, Foster, Dean, Kakade, Sham |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
A New Perspective on Shampoo's Preconditioner
von: Morwani, Depen, et al.
Veröffentlicht: (2024)
von: Morwani, Depen, et al.
Veröffentlicht: (2024)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
von: Morwani, Depen, et al.
Veröffentlicht: (2025)
von: Morwani, Depen, et al.
Veröffentlicht: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
von: Abreu, Natalie, et al.
Veröffentlicht: (2025)
von: Abreu, Natalie, et al.
Veröffentlicht: (2025)
Sample-Efficient Agnostic Boosting
von: Ghai, Udaya, et al.
Veröffentlicht: (2024)
von: Ghai, Udaya, et al.
Veröffentlicht: (2024)
Deconstructing What Makes a Good Optimizer for Language Models
von: Zhao, Rosie, et al.
Veröffentlicht: (2024)
von: Zhao, Rosie, et al.
Veröffentlicht: (2024)
Neural Coordination and Capacity Control for Inventory Management
von: Eisenach, Carson, et al.
Veröffentlicht: (2024)
von: Eisenach, Carson, et al.
Veröffentlicht: (2024)
Beyond Implicit Bias: The Insignificance of SGD Noise in Online Learning
von: Vyas, Nikhil, et al.
Veröffentlicht: (2023)
von: Vyas, Nikhil, et al.
Veröffentlicht: (2023)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
von: Song, Yuda, et al.
Veröffentlicht: (2024)
von: Song, Yuda, et al.
Veröffentlicht: (2024)
A Mechanism Study of Delayed Loss Spikes in Batch-Normalized Linear Models
von: Gao, Peifeng, et al.
Veröffentlicht: (2026)
von: Gao, Peifeng, et al.
Veröffentlicht: (2026)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
von: Kou, Yiwen, et al.
Veröffentlicht: (2024)
von: Kou, Yiwen, et al.
Veröffentlicht: (2024)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
von: Liu, Bingbin, et al.
Veröffentlicht: (2025)
von: Liu, Bingbin, et al.
Veröffentlicht: (2025)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
von: Ostroukhov, Petr, et al.
Veröffentlicht: (2024)
von: Ostroukhov, Petr, et al.
Veröffentlicht: (2024)
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
von: Wang, Jinbo, et al.
Veröffentlicht: (2026)
von: Wang, Jinbo, et al.
Veröffentlicht: (2026)
On the Batch Size Selection in Stochastic Gradient Methods Using No-Replacement Sampling
von: Boresta, Marco, et al.
Veröffentlicht: (2025)
von: Boresta, Marco, et al.
Veröffentlicht: (2025)
SOAP: Improving and Stabilizing Shampoo using Adam
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
LOTION: Smoothing the Optimization Landscape for Quantized Training
von: Kwun, Mujin, et al.
Veröffentlicht: (2025)
von: Kwun, Mujin, et al.
Veröffentlicht: (2025)
Joint Order Selection, Allocation, Batching and Picking for Large Scale Warehouses
von: Abelli, Giorgio, et al.
Veröffentlicht: (2024)
von: Abelli, Giorgio, et al.
Veröffentlicht: (2024)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
Beyond First-Order Methods for $\ell_p$-Structured Non-Monotone Variational Inequalities
von: Vyas, Abhijeet, et al.
Veröffentlicht: (2026)
von: Vyas, Abhijeet, et al.
Veröffentlicht: (2026)
Mirror-Free Proximal Methods
von: Vyas, Abhijeet, et al.
Veröffentlicht: (2026)
von: Vyas, Abhijeet, et al.
Veröffentlicht: (2026)
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
Constraint Programming Models For Serial Batch Scheduling With Minimum Batch Size
von: Huertas, Jorge A., et al.
Veröffentlicht: (2025)
von: Huertas, Jorge A., et al.
Veröffentlicht: (2025)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
Robustness of Iteratively Pre-Conditioned Gradient-Descent Method: The Case of Distributed Linear Regression Problem
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2021)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2021)
Iterative Pre-Conditioning for Expediting the Gradient-Descent Method: The Distributed Linear Least-Squares Problem
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2020)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2020)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
von: Imaizumi, Kento, et al.
Veröffentlicht: (2025)
von: Imaizumi, Kento, et al.
Veröffentlicht: (2025)
An Improved Analysis of Langevin Algorithms with Prior Diffusion for Non-Log-Concave Sampling
von: Huang, Xunpeng, et al.
Veröffentlicht: (2024)
von: Huang, Xunpeng, et al.
Veröffentlicht: (2024)
Dynamic Batching of Online Arrivals to Leverage Economies of Scale
von: Bhimaraju, Akhil, et al.
Veröffentlicht: (2023)
von: Bhimaraju, Akhil, et al.
Veröffentlicht: (2023)
Faster Sampling without Isoperimetry via Diffusion-based Monte Carlo
von: Huang, Xunpeng, et al.
Veröffentlicht: (2024)
von: Huang, Xunpeng, et al.
Veröffentlicht: (2024)
Feature emergence via margin maximization: case studies in algebraic tasks
von: Morwani, Depen, et al.
Veröffentlicht: (2023)
von: Morwani, Depen, et al.
Veröffentlicht: (2023)
Convergence of Riemannian Stochastic Gradient Descents: Varying Batch Sizes And Nonstandard Batch Forming
von: Wu, Hao
Veröffentlicht: (2026)
von: Wu, Hao
Veröffentlicht: (2026)
Quasi-Newton Compatible Actor-Critic for Deterministic Policies
von: Kordabad, Arash Bahari, et al.
Veröffentlicht: (2025)
von: Kordabad, Arash Bahari, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025) -
A New Perspective on Shampoo's Preconditioner
von: Morwani, Depen, et al.
Veröffentlicht: (2024) -
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025) -
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
von: Morwani, Depen, et al.
Veröffentlicht: (2025) -
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)