Marek, M., Lotfi, S., Somasundaram, A., Wilson, A. G., & Goldblum, M. (2025). Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful.
Citazione stile Chigago Style (17a edizione)Marek, Martin, Sanae Lotfi, Aditya Somasundaram, Andrew Gordon Wilson, e Micah Goldblum. Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful. 2025.
Citatione MLA (9a ed.)Marek, Martin, et al. Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful. 2025.
Attenzione: Queste citazioni potrebbero non essere precise al 100%.