Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Merrill, William, Arora, Shane, Groeneveld, Dirk, Hajishirzi, Hannaneh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
Scaling Law for Language Models Training Considering Batch Size
von: Shuai, Xian, et al.
Veröffentlicht: (2024)
von: Shuai, Xian, et al.
Veröffentlicht: (2024)
Convergence Bound and Critical Batch Size of Muon Optimizer
von: Sato, Naoki, et al.
Veröffentlicht: (2025)
von: Sato, Naoki, et al.
Veröffentlicht: (2025)
Revisiting LARS for Large Batch Training Generalization of Neural Networks
von: Do, Khoi, et al.
Veröffentlicht: (2023)
von: Do, Khoi, et al.
Veröffentlicht: (2023)
Full-Graph vs. Mini-Batch Training: Comprehensive Analysis from a Batch Size and Fan-Out Size Perspective
von: Liu, Mengfan, et al.
Veröffentlicht: (2026)
von: Liu, Mengfan, et al.
Veröffentlicht: (2026)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
von: Ostroukhov, Petr, et al.
Veröffentlicht: (2024)
von: Ostroukhov, Petr, et al.
Veröffentlicht: (2024)
A Simple and Efficient Approach to Batch Bayesian Optimization
von: Zhan, Dawei, et al.
Veröffentlicht: (2024)
von: Zhan, Dawei, et al.
Veröffentlicht: (2024)
How Does Critical Batch Size Scale in Pre-training?
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
DIVEBATCH: Accelerating Model Training Through Gradient-Diversity Aware Batch Size Adaptation
von: Chen, Yuen, et al.
Veröffentlicht: (2025)
von: Chen, Yuen, et al.
Veröffentlicht: (2025)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
von: Marek, Martin, et al.
Veröffentlicht: (2025)
von: Marek, Martin, et al.
Veröffentlicht: (2025)
Learning to Detect Language Model Training Data via Active Reconstruction
von: Yin, Junjie Oscar, et al.
Veröffentlicht: (2026)
von: Yin, Junjie Oscar, et al.
Veröffentlicht: (2026)
Stochastic Normalized Gradient Descent with Momentum for Large-Batch Training
von: Zhao, Shen-Yi, et al.
Veröffentlicht: (2020)
von: Zhao, Shen-Yi, et al.
Veröffentlicht: (2020)
Enabling Large Batch Size Training for DNN Models Beyond the Memory Limit While Maintaining Performance
von: Piao, XinYu, et al.
Veröffentlicht: (2021)
von: Piao, XinYu, et al.
Veröffentlicht: (2021)
Diversified Batch Selection for Training Acceleration
von: Hong, Feng, et al.
Veröffentlicht: (2024)
von: Hong, Feng, et al.
Veröffentlicht: (2024)
EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
von: Yoon, Junsang, et al.
Veröffentlicht: (2024)
von: Yoon, Junsang, et al.
Veröffentlicht: (2024)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
Smaller Batches, Bigger Gains? Investigating the Impact of Batch Sizes on Reinforcement Learning Based Real-World Production Scheduling
von: Müller, Arthur, et al.
Veröffentlicht: (2024)
von: Müller, Arthur, et al.
Veröffentlicht: (2024)
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
von: Kim, Joongwon, et al.
Veröffentlicht: (2024)
von: Kim, Joongwon, et al.
Veröffentlicht: (2024)
Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling
von: Li, Shuaipeng, et al.
Veröffentlicht: (2024)
von: Li, Shuaipeng, et al.
Veröffentlicht: (2024)
Collaborative Batch Size Optimization for Federated Learning
von: Geimer, Arno, et al.
Veröffentlicht: (2025)
von: Geimer, Arno, et al.
Veröffentlicht: (2025)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
von: Chen, Tong, et al.
Veröffentlicht: (2025)
von: Chen, Tong, et al.
Veröffentlicht: (2025)
An Investigation of Batch Normalization in Off-Policy Actor-Critic Algorithms
von: Wang, Li, et al.
Veröffentlicht: (2025)
von: Wang, Li, et al.
Veröffentlicht: (2025)
Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training
von: Tyagi, Sahil, et al.
Veröffentlicht: (2026)
von: Tyagi, Sahil, et al.
Veröffentlicht: (2026)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025)
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025)
Riemannian Batch Normalization: A Gyro Approach
von: Chen, Ziheng, et al.
Veröffentlicht: (2025)
von: Chen, Ziheng, et al.
Veröffentlicht: (2025)
Subsampling is not Magic: Why Large Batch Sizes Work for Differentially Private Stochastic Optimisation
von: Räisä, Ossi, et al.
Veröffentlicht: (2024)
von: Räisä, Ossi, et al.
Veröffentlicht: (2024)
Beware of the Batch Size: Hyperparameter Bias in Evaluating LoRA
von: Lee, Sangyoon, et al.
Veröffentlicht: (2026)
von: Lee, Sangyoon, et al.
Veröffentlicht: (2026)
Robust Batched Bandits
von: Guo, Yunwen, et al.
Veröffentlicht: (2025)
von: Guo, Yunwen, et al.
Veröffentlicht: (2025)
Supervised Batch Normalization
von: Faye, Bilal, et al.
Veröffentlicht: (2024)
von: Faye, Bilal, et al.
Veröffentlicht: (2024)
Batch Bayesian Active Learning with Partial Batch Label Sampling
von: Hu, Kangping, et al.
Veröffentlicht: (2025)
von: Hu, Kangping, et al.
Veröffentlicht: (2025)
Adaptive Batch Sizes for Active Learning A Probabilistic Numerics Approach
von: Adachi, Masaki, et al.
Veröffentlicht: (2023)
von: Adachi, Masaki, et al.
Veröffentlicht: (2023)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
Edge Intelligence Optimization for Large Language Model Inference with Batching and Quantization
von: Zhang, Xinyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xinyuan, et al.
Veröffentlicht: (2024)
Decoding-Time Language Model Alignment with Multiple Objectives
von: Shi, Ruizhe, et al.
Veröffentlicht: (2024)
von: Shi, Ruizhe, et al.
Veröffentlicht: (2024)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
Merge to Learn: Efficiently Adding Skills to Language Models with Model Merging
von: Morrison, Jacob, et al.
Veröffentlicht: (2024)
von: Morrison, Jacob, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
von: Zhao, Bowen, et al.
Veröffentlicht: (2024) -
Scaling Law for Language Models Training Considering Batch Size
von: Shuai, Xian, et al.
Veröffentlicht: (2024) -
Convergence Bound and Critical Batch Size of Muon Optimizer
von: Sato, Naoki, et al.
Veröffentlicht: (2025) -
Revisiting LARS for Large Batch Training Generalization of Neural Networks
von: Do, Khoi, et al.
Veröffentlicht: (2023) -
Full-Graph vs. Mini-Batch Training: Comprehensive Analysis from a Batch Size and Fan-Out Size Perspective
von: Liu, Mengfan, et al.
Veröffentlicht: (2026)