Analytic theory of dropout regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mori, Francesco, Mignacco, Francesca
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908526224343040
author Mori, Francesco
Mignacco, Francesca
author_facet Mori, Francesco
Mignacco, Francesca
contents Dropout is a regularization technique widely used in training artificial neural networks to mitigate overfitting. It consists of dynamically deactivating subsets of the network during training to promote more robust representations. Despite its widespread adoption, dropout probabilities are often selected heuristically, and theoretical explanations of its success remain sparse. Here, we analytically study dropout in two-layer neural networks trained with online stochastic gradient descent. In the high-dimensional limit, we derive a set of ordinary differential equations that fully characterize the evolution of the network during training and capture the effects of dropout. We obtain a number of exact results describing the generalization error and the optimal dropout probability at short, intermediate, and long training times. Our analysis shows that dropout reduces detrimental correlations between hidden nodes, mitigates the impact of label noise, and that the optimal dropout probability increases with the level of noise in the data. Our results are validated by extensive numerical simulations.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07792
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Analytic theory of dropout regularization
Mori, Francesco
Mignacco, Francesca
Machine Learning
Disordered Systems and Neural Networks
Statistical Mechanics
Dropout is a regularization technique widely used in training artificial neural networks to mitigate overfitting. It consists of dynamically deactivating subsets of the network during training to promote more robust representations. Despite its widespread adoption, dropout probabilities are often selected heuristically, and theoretical explanations of its success remain sparse. Here, we analytically study dropout in two-layer neural networks trained with online stochastic gradient descent. In the high-dimensional limit, we derive a set of ordinary differential equations that fully characterize the evolution of the network during training and capture the effects of dropout. We obtain a number of exact results describing the generalization error and the optimal dropout probability at short, intermediate, and long training times. Our analysis shows that dropout reduces detrimental correlations between hidden nodes, mitigates the impact of label noise, and that the optimal dropout probability increases with the level of noise in the data. Our results are validated by extensive numerical simulations.
title Analytic theory of dropout regularization
topic Machine Learning
Disordered Systems and Neural Networks
Statistical Mechanics
url https://arxiv.org/abs/2505.07792