When Bias Meets Trainability: Connecting Theories of Initialization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bassi, Alberto, Baity-Jesi, Marco, Lucchi, Aurelien, Albert, Carlo, Francazi, Emanuele
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918362063306752
author Bassi, Alberto
Baity-Jesi, Marco
Lucchi, Aurelien
Albert, Carlo
Francazi, Emanuele
author_facet Bassi, Alberto
Baity-Jesi, Marco
Lucchi, Aurelien
Albert, Carlo
Francazi, Emanuele
contents The statistical properties of deep neural networks (DNNs) at initialization play an important role to comprehend their trainability and the intrinsic architectural biases they possess before data exposure Well established mean field (MF) theories have uncovered that the distribution of parameters of randomly initialized networks strongly influences the behavior of the gradients, dictating whether they explode or vanish. Recent work has showed that untrained DNNs also manifest an initial guessing bias (IGB), in which large regions of the input space are assigned to a single class. In this work, we provide a theoretical proof that links IGB to previous MF theories for a vast class of DNNs, showing that efficient learning is tightly connected to a network prejudice towards a specific class. This connection leads to a counterintuitive conclusion: the initialization that optimizes trainability is systematically biased rather than neutral.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12096
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Bias Meets Trainability: Connecting Theories of Initialization
Bassi, Alberto
Baity-Jesi, Marco
Lucchi, Aurelien
Albert, Carlo
Francazi, Emanuele
Machine Learning
Artificial Intelligence
The statistical properties of deep neural networks (DNNs) at initialization play an important role to comprehend their trainability and the intrinsic architectural biases they possess before data exposure Well established mean field (MF) theories have uncovered that the distribution of parameters of randomly initialized networks strongly influences the behavior of the gradients, dictating whether they explode or vanish. Recent work has showed that untrained DNNs also manifest an initial guessing bias (IGB), in which large regions of the input space are assigned to a single class. In this work, we provide a theoretical proof that links IGB to previous MF theories for a vast class of DNNs, showing that efficient learning is tightly connected to a network prejudice towards a specific class. This connection leads to a counterintuitive conclusion: the initialization that optimizes trainability is systematically biased rather than neutral.
title When Bias Meets Trainability: Connecting Theories of Initialization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.12096