Opening the Black Box: predicting the trainability of deep neural networks with reconstruction entropy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thurn, Yanick, Jefferson, Ro, Erdmenger, Johanna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912162391261184
author Thurn, Yanick
Jefferson, Ro
Erdmenger, Johanna
author_facet Thurn, Yanick
Jefferson, Ro
Erdmenger, Johanna
contents An important challenge in machine learning is to predict the initial conditions under which a given neural network will be trainable. We present a method for predicting the trainable regime in parameter space for deep feedforward neural networks (DNNs) based on reconstructing the input from subsequent activation layers via a cascade of single-layer auxiliary networks. We show that a single epoch of training of the shallow cascade networks is sufficient to predict the trainability of the deep feedforward network on a range of datasets (MNIST, CIFAR10, FashionMNIST, and white noise), thereby providing a significant reduction in overall training time. We achieve this by computing the relative entropy between reconstructed images and the original inputs, and show that this probe of information loss is sensitive to the phase behaviour of the network. We further demonstrate that this method generalizes to residual neural networks (ResNets) and convolutional neural networks (CNNs). Moreover, our method illustrates the network's decision making process by displaying the changes performed on the input data at each layer, which we demonstrate for both a DNN trained on MNIST and the vgg16 CNN trained on the ImageNet dataset. Our results provide a technique for significantly accelerating the training of large neural networks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12916
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Opening the Black Box: predicting the trainability of deep neural networks with reconstruction entropy
Thurn, Yanick
Jefferson, Ro
Erdmenger, Johanna
Machine Learning
Disordered Systems and Neural Networks
High Energy Physics - Theory
An important challenge in machine learning is to predict the initial conditions under which a given neural network will be trainable. We present a method for predicting the trainable regime in parameter space for deep feedforward neural networks (DNNs) based on reconstructing the input from subsequent activation layers via a cascade of single-layer auxiliary networks. We show that a single epoch of training of the shallow cascade networks is sufficient to predict the trainability of the deep feedforward network on a range of datasets (MNIST, CIFAR10, FashionMNIST, and white noise), thereby providing a significant reduction in overall training time. We achieve this by computing the relative entropy between reconstructed images and the original inputs, and show that this probe of information loss is sensitive to the phase behaviour of the network. We further demonstrate that this method generalizes to residual neural networks (ResNets) and convolutional neural networks (CNNs). Moreover, our method illustrates the network's decision making process by displaying the changes performed on the input data at each layer, which we demonstrate for both a DNN trained on MNIST and the vgg16 CNN trained on the ImageNet dataset. Our results provide a technique for significantly accelerating the training of large neural networks.
title Opening the Black Box: predicting the trainability of deep neural networks with reconstruction entropy
topic Machine Learning
Disordered Systems and Neural Networks
High Energy Physics - Theory
url https://arxiv.org/abs/2406.12916