DIVEBATCH: Accelerating Model Training Through Gradient-Diversity Aware Batch Size Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yuen, Wang, Yian, Sundaram, Hari |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
State Contamination in Memory-Augmented LLM Agents
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
On the Necessity of Output Distribution Reweighting for Effective Class Unlearning
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2025)
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2025)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024)
by: Umeda, Hikaru, et al.
Published: (2024)
Diversified Batch Selection for Training Acceleration
by: Hong, Feng, et al.
Published: (2024)
by: Hong, Feng, et al.
Published: (2024)
LOTOS: Layer-wise Orthogonalization for Training Robust Ensembles
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2024)
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2024)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
by: Marek, Martin, et al.
Published: (2025)
by: Marek, Martin, et al.
Published: (2025)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
Scaling Law for Language Models Training Considering Batch Size
by: Shuai, Xian, et al.
Published: (2024)
by: Shuai, Xian, et al.
Published: (2024)
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
by: Merrill, William, et al.
Published: (2025)
by: Merrill, William, et al.
Published: (2025)
Beyond Batch Learning: Global Awareness Enhanced Domain Adaptation
by: Luo, Lingkun, et al.
Published: (2025)
by: Luo, Lingkun, et al.
Published: (2025)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Efficient GNN Training Through Structure-Aware Randomized Mini-Batching
by: Balaji, Vignesh, et al.
Published: (2025)
by: Balaji, Vignesh, et al.
Published: (2025)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
by: Kamo, Keisuke, et al.
Published: (2025)
by: Kamo, Keisuke, et al.
Published: (2025)
Full-Graph vs. Mini-Batch Training: Comprehensive Analysis from a Batch Size and Fan-Out Size Perspective
by: Liu, Mengfan, et al.
Published: (2026)
by: Liu, Mengfan, et al.
Published: (2026)
CEV-LM: Controlled Edit Vector Language Model for Shaping Natural Language Generations
by: Moorjani, Samraj, et al.
Published: (2024)
by: Moorjani, Samraj, et al.
Published: (2024)
Stochastic Normalized Gradient Descent with Momentum for Large-Batch Training
by: Zhao, Shen-Yi, et al.
Published: (2020)
by: Zhao, Shen-Yi, et al.
Published: (2020)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
by: Oowada, Kanata, et al.
Published: (2025)
by: Oowada, Kanata, et al.
Published: (2025)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
by: Kondo, Yuichi, et al.
Published: (2025)
by: Kondo, Yuichi, et al.
Published: (2025)
Do You Trust the Process?: Modeling Institutional Trust for Community Adoption of Reinforcement Learning Policies
by: Balepur, Naina, et al.
Published: (2025)
by: Balepur, Naina, et al.
Published: (2025)
Component-Aware Pruning Framework for Neural Network Controllers via Gradient-Based Importance Estimation
by: Sundaram, Ganesh, et al.
Published: (2026)
by: Sundaram, Ganesh, et al.
Published: (2026)
One Size Does Not Fit All: Architecture-Aware Adaptive Batch Scheduling with DEBA
by: Belias, François, et al.
Published: (2025)
by: Belias, François, et al.
Published: (2025)
AMUN: Adversarial Machine UNlearning
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2025)
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2025)
Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling
by: Li, Shuaipeng, et al.
Published: (2024)
by: Li, Shuaipeng, et al.
Published: (2024)
Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
by: Zou, Heming, et al.
Published: (2025)
by: Zou, Heming, et al.
Published: (2025)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
Convergence Bound and Critical Batch Size of Muon Optimizer
by: Sato, Naoki, et al.
Published: (2025)
by: Sato, Naoki, et al.
Published: (2025)
Batched Low-Rank Adaptation of Foundation Models
by: Wen, Yeming, et al.
Published: (2023)
by: Wen, Yeming, et al.
Published: (2023)
Is BatchEnsemble a Single Model? On Calibration and Diversity of Efficient Ensembles
by: Zamyatin, Anton, et al.
Published: (2026)
by: Zamyatin, Anton, et al.
Published: (2026)
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
by: Lu, Yishun, et al.
Published: (2025)
by: Lu, Yishun, et al.
Published: (2025)
Training Data Selection with Gradient Orthogonality for Efficient Domain Adaptation
by: Zhang, Xiyang, et al.
Published: (2026)
by: Zhang, Xiyang, et al.
Published: (2026)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
by: Harada, Hinata, et al.
Published: (2024)
by: Harada, Hinata, et al.
Published: (2024)
Enabling Large Batch Size Training for DNN Models Beyond the Memory Limit While Maintaining Performance
by: Piao, XinYu, et al.
Published: (2021)
by: Piao, XinYu, et al.
Published: (2021)
Spectrum Extraction and Clipping for Implicitly Linear Layers
by: Boroojeny, Ali Ebrahimpour, et al.
Published: (2024)
by: Boroojeny, Ali Ebrahimpour, et al.
Published: (2024)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
by: Tabesh, Soroush, et al.
Published: (2025)
by: Tabesh, Soroush, et al.
Published: (2025)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
Similar Items
-
State Contamination in Memory-Augmented LLM Agents
by: Wang, Yian, et al.
Published: (2026) -
On the Necessity of Output Distribution Reweighting for Effective Class Unlearning
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2025) -
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024) -
Diversified Batch Selection for Training Acceleration
by: Hong, Feng, et al.
Published: (2024) -
LOTOS: Layer-wise Orthogonalization for Training Robust Ensembles
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2024)