Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
Fuente:
arXiv
Guardado en:
| Autores principales: | Marek, Martin, Lotfi, Sanae, Somasundaram, Aditya, Wilson, Andrew Gordon, Goldblum, Micah |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
por: Lotfi, Sanae, et al.
Publicado: (2024)
por: Lotfi, Sanae, et al.
Publicado: (2024)
Non-Vacuous Generalization Bounds for Large Language Models
por: Lotfi, Sanae, et al.
Publicado: (2023)
por: Lotfi, Sanae, et al.
Publicado: (2023)
The Lie Derivative for Measuring Learned Equivariance
por: Gruver, Nate, et al.
Publicado: (2022)
por: Gruver, Nate, et al.
Publicado: (2022)
The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning
por: Goldblum, Micah, et al.
Publicado: (2023)
por: Goldblum, Micah, et al.
Publicado: (2023)
Compute Better Spent: Replacing Dense Layers with Structured Matrices
por: Qiu, Shikai, et al.
Publicado: (2024)
por: Qiu, Shikai, et al.
Publicado: (2024)
Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes
por: Jayawardhana, Mayuka, et al.
Publicado: (2025)
por: Jayawardhana, Mayuka, et al.
Publicado: (2025)
Closing the Train-Test Gap in World Models for Gradient-Based Planning
por: Parthasarathy, Arjun, et al.
Publicado: (2025)
por: Parthasarathy, Arjun, et al.
Publicado: (2025)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
por: Kuang, Yilun, et al.
Publicado: (2025)
por: Kuang, Yilun, et al.
Publicado: (2025)
Uncertainty Drives Social Bias Changes in Quantized Large Language Models
por: Hua, Stanley Z., et al.
Publicado: (2026)
por: Hua, Stanley Z., et al.
Publicado: (2026)
Large Language Models Must Be Taught to Know What They Don't Know
por: Kapoor, Sanyam, et al.
Publicado: (2024)
por: Kapoor, Sanyam, et al.
Publicado: (2024)
Adaptive Retention & Correction: Test-Time Training for Continual Learning
por: Chen, Haoran, et al.
Publicado: (2024)
por: Chen, Haoran, et al.
Publicado: (2024)
Just How Flexible are Neural Networks in Practice?
por: Shwartz-Ziv, Ravid, et al.
Publicado: (2024)
por: Shwartz-Ziv, Ravid, et al.
Publicado: (2024)
Subsampling is not Magic: Why Large Batch Sizes Work for Differentially Private Stochastic Optimisation
por: Räisä, Ossi, et al.
Publicado: (2024)
por: Räisä, Ossi, et al.
Publicado: (2024)
Scaling Law for Language Models Training Considering Batch Size
por: Shuai, Xian, et al.
Publicado: (2024)
por: Shuai, Xian, et al.
Publicado: (2024)
Why Masking Diffusion Works: Condition on the Jump Schedule for Improved Discrete Diffusion
por: Amin, Alan N., et al.
Publicado: (2025)
por: Amin, Alan N., et al.
Publicado: (2025)
DIVEBATCH: Accelerating Model Training Through Gradient-Diversity Aware Batch Size Adaptation
por: Chen, Yuen, et al.
Publicado: (2025)
por: Chen, Yuen, et al.
Publicado: (2025)
Efficient Conditioning Why Pseudo Observation Batch Bayesian Optimization Works When It Does not
por: Nagaswetha, Kumbha, et al.
Publicado: (2026)
por: Nagaswetha, Kumbha, et al.
Publicado: (2026)
Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
por: Pal, Arka, et al.
Publicado: (2025)
por: Pal, Arka, et al.
Publicado: (2025)
On Training in Imagination
por: Timor, Nadav, et al.
Publicado: (2026)
por: Timor, Nadav, et al.
Publicado: (2026)
Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
por: Lotfi, Sanae, et al.
Publicado: (2026)
por: Lotfi, Sanae, et al.
Publicado: (2026)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
por: Umeda, Hikaru, et al.
Publicado: (2025)
por: Umeda, Hikaru, et al.
Publicado: (2025)
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
por: Merrill, William, et al.
Publicado: (2025)
por: Merrill, William, et al.
Publicado: (2025)
Enhancing SignSGD: Small-Batch Convergence Analysis and a Hybrid Switching Strategy
por: Chen, Haoran, et al.
Publicado: (2026)
por: Chen, Haoran, et al.
Publicado: (2026)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
por: Jain, Neel, et al.
Publicado: (2024)
por: Jain, Neel, et al.
Publicado: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
por: Hayes, Kevin David, et al.
Publicado: (2025)
por: Hayes, Kevin David, et al.
Publicado: (2025)
vTune: Verifiable Fine-Tuning for LLMs Through Backdooring
por: Zhang, Eva, et al.
Publicado: (2024)
por: Zhang, Eva, et al.
Publicado: (2024)
Dynamic Delayed Tree Expansion For Improved Multi-Path Speculative Decoding
por: Thomas, Rahul, et al.
Publicado: (2026)
por: Thomas, Rahul, et al.
Publicado: (2026)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
por: Marek, Martin, et al.
Publicado: (2026)
por: Marek, Martin, et al.
Publicado: (2026)
Asynchronous Local-SGD Training for Language Modeling
por: Liu, Bo, et al.
Publicado: (2024)
por: Liu, Bo, et al.
Publicado: (2024)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
por: Tao, Hongyi, et al.
Publicado: (2026)
por: Tao, Hongyi, et al.
Publicado: (2026)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
por: Islamov, Rustem, et al.
Publicado: (2026)
por: Islamov, Rustem, et al.
Publicado: (2026)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
por: Sakip, Akhmed, et al.
Publicado: (2026)
por: Sakip, Akhmed, et al.
Publicado: (2026)
Convergence of Riemannian Stochastic Gradient Descents: Varying Batch Sizes And Nonstandard Batch Forming
por: Wu, Hao
Publicado: (2026)
por: Wu, Hao
Publicado: (2026)
Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices
por: Potapczynski, Andres, et al.
Publicado: (2024)
por: Potapczynski, Andres, et al.
Publicado: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
por: Ostroukhov, Petr, et al.
Publicado: (2024)
por: Ostroukhov, Petr, et al.
Publicado: (2024)
On the Utility of Equal Batch Sizes for Inference in Stochastic Gradient Descent
por: Singh, Rahul, et al.
Publicado: (2023)
por: Singh, Rahul, et al.
Publicado: (2023)
Accumulative SGD Influence Estimation for Data Attribution
por: Shi, Yunxiao, et al.
Publicado: (2025)
por: Shi, Yunxiao, et al.
Publicado: (2025)
Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion
por: Souri, Hossein, et al.
Publicado: (2024)
por: Souri, Hossein, et al.
Publicado: (2024)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
por: Kovačević, Filip, et al.
Publicado: (2026)
por: Kovačević, Filip, et al.
Publicado: (2026)
Ejemplares similares
-
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
por: Lotfi, Sanae, et al.
Publicado: (2024) -
Non-Vacuous Generalization Bounds for Large Language Models
por: Lotfi, Sanae, et al.
Publicado: (2023) -
The Lie Derivative for Measuring Learned Equivariance
por: Gruver, Nate, et al.
Publicado: (2022) -
The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning
por: Goldblum, Micah, et al.
Publicado: (2023) -
Compute Better Spent: Replacing Dense Layers with Structured Matrices
por: Qiu, Shikai, et al.
Publicado: (2024)