Phase Diagram of Initial Condensation for Two-layer Neural Networks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Zhengan, Li, Yuqing, Luo, Tao, Zhou, Zhangchen, Xu, Zhi-Qin John
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918421793341440
author Chen, Zhengan
Li, Yuqing
Luo, Tao
Zhou, Zhangchen
Xu, Zhi-Qin John
author_facet Chen, Zhengan
Li, Yuqing
Luo, Tao
Zhou, Zhangchen
Xu, Zhi-Qin John
contents The phenomenon of distinct behaviors exhibited by neural networks under varying scales of initialization remains an enigma in deep learning research. In this paper, based on the earlier work by Luo et al.~\cite{luo2021phase}, we present a phase diagram of initial condensation for two-layer neural networks. Condensation is a phenomenon wherein the weight vectors of neural networks concentrate on isolated orientations during the training process, and it is a feature in non-linear learning process that enables neural networks to possess better generalization abilities. Our phase diagram serves to provide a comprehensive understanding of the dynamical regimes of neural networks and their dependence on the choice of hyperparameters related to initialization. Furthermore, we demonstrate in detail the underlying mechanisms by which small initialization leads to condensation at the initial training stage.
format Preprint
id arxiv_https___arxiv_org_abs_2303_06561
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Phase Diagram of Initial Condensation for Two-layer Neural Networks
Chen, Zhengan
Li, Yuqing
Luo, Tao
Zhou, Zhangchen
Xu, Zhi-Qin John
Machine Learning
Disordered Systems and Neural Networks
Optimization and Control
68U99, 90C26, 34A45
The phenomenon of distinct behaviors exhibited by neural networks under varying scales of initialization remains an enigma in deep learning research. In this paper, based on the earlier work by Luo et al.~\cite{luo2021phase}, we present a phase diagram of initial condensation for two-layer neural networks. Condensation is a phenomenon wherein the weight vectors of neural networks concentrate on isolated orientations during the training process, and it is a feature in non-linear learning process that enables neural networks to possess better generalization abilities. Our phase diagram serves to provide a comprehensive understanding of the dynamical regimes of neural networks and their dependence on the choice of hyperparameters related to initialization. Furthermore, we demonstrate in detail the underlying mechanisms by which small initialization leads to condensation at the initial training stage.
title Phase Diagram of Initial Condensation for Two-layer Neural Networks
topic Machine Learning
Disordered Systems and Neural Networks
Optimization and Control
68U99, 90C26, 34A45
url https://arxiv.org/abs/2303.06561