Phase transitions in the mini-batch size for sparse and dense two-layer neural networks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Marino, Raffaele, Ricci-Tersenghi, Federico
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913194084139008
author Marino, Raffaele
Ricci-Tersenghi, Federico
author_facet Marino, Raffaele
Ricci-Tersenghi, Federico
contents The use of mini-batches of data in training artificial neural networks is nowadays very common. Despite its broad usage, theories explaining quantitatively how large or small the optimal mini-batch size should be are missing. This work presents a systematic attempt at understanding the role of the mini-batch size in training two-layer neural networks. Working in the teacher-student scenario, with a sparse teacher, and focusing on tasks of different complexity, we quantify the effects of changing the mini-batch size $m$. We find that often the generalization performances of the student strongly depend on $m$ and may undergo sharp phase transitions at a critical value $m_c$, such that for $m<m_c$ the training process fails, while for $m>m_c$ the student learns perfectly or generalizes very well the teacher. Phase transitions are induced by collective phenomena firstly discovered in statistical mechanics and later observed in many fields of science. Observing a phase transition by varying the mini-batch size across different architectures raises several questions about the role of this hyperparameter in the neural network learning process.
format Preprint
id arxiv_https___arxiv_org_abs_2305_06435
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Phase transitions in the mini-batch size for sparse and dense two-layer neural networks
Marino, Raffaele
Ricci-Tersenghi, Federico
Disordered Systems and Neural Networks
Statistical Mechanics
Artificial Intelligence
Machine Learning
The use of mini-batches of data in training artificial neural networks is nowadays very common. Despite its broad usage, theories explaining quantitatively how large or small the optimal mini-batch size should be are missing. This work presents a systematic attempt at understanding the role of the mini-batch size in training two-layer neural networks. Working in the teacher-student scenario, with a sparse teacher, and focusing on tasks of different complexity, we quantify the effects of changing the mini-batch size $m$. We find that often the generalization performances of the student strongly depend on $m$ and may undergo sharp phase transitions at a critical value $m_c$, such that for $m<m_c$ the training process fails, while for $m>m_c$ the student learns perfectly or generalizes very well the teacher. Phase transitions are induced by collective phenomena firstly discovered in statistical mechanics and later observed in many fields of science. Observing a phase transition by varying the mini-batch size across different architectures raises several questions about the role of this hyperparameter in the neural network learning process.
title Phase transitions in the mini-batch size for sparse and dense two-layer neural networks
topic Disordered Systems and Neural Networks
Statistical Mechanics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2305.06435