Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Min, Hancheng, Mallada, Enrique, Vidal, René
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916173893861376
author Min, Hancheng
Mallada, Enrique
Vidal, René
author_facet Min, Hancheng
Mallada, Enrique
Vidal, René
contents This paper studies the problem of training a two-layer ReLU network for binary classification using gradient flow with small initialization. We consider a training dataset with well-separated input vectors: Any pair of input data with the same label are positively correlated, and any pair with different labels are negatively correlated. Our analysis shows that, during the early phase of training, neurons in the first layer try to align with either the positive data or the negative data, depending on its corresponding weight on the second layer. A careful analysis of the neurons' directional dynamics allows us to provide an $\mathcal{O}(\frac{\log n}{\sqrtμ})$ upper bound on the time it takes for all neurons to achieve good alignment with the input data, where $n$ is the number of data points and $μ$ measures how well the data are separated. After the early alignment phase, the loss converges to zero at a $\mathcal{O}(\frac{1}{t})$ rate, and the weight matrix on the first layer is approximately low-rank. Numerical experiments on the MNIST dataset illustrate our theoretical findings.
format Preprint
id arxiv_https___arxiv_org_abs_2307_12851
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization
Min, Hancheng
Mallada, Enrique
Vidal, René
Machine Learning
This paper studies the problem of training a two-layer ReLU network for binary classification using gradient flow with small initialization. We consider a training dataset with well-separated input vectors: Any pair of input data with the same label are positively correlated, and any pair with different labels are negatively correlated. Our analysis shows that, during the early phase of training, neurons in the first layer try to align with either the positive data or the negative data, depending on its corresponding weight on the second layer. A careful analysis of the neurons' directional dynamics allows us to provide an $\mathcal{O}(\frac{\log n}{\sqrtμ})$ upper bound on the time it takes for all neurons to achieve good alignment with the input data, where $n$ is the number of data points and $μ$ measures how well the data are separated. After the early alignment phase, the loss converges to zero at a $\mathcal{O}(\frac{1}{t})$ rate, and the weight matrix on the first layer is approximately low-rank. Numerical experiments on the MNIST dataset illustrate our theoretical findings.
title Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization
topic Machine Learning
url https://arxiv.org/abs/2307.12851