A topological description of loss surfaces based on Betti Numbers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bucarelli, Maria Sofia, D'Inverno, Giuseppe Alessio, Bianchini, Monica, Scarselli, Franco, Silvestri, Fabrizio
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909064937603072
author Bucarelli, Maria Sofia
D'Inverno, Giuseppe Alessio
Bianchini, Monica
Scarselli, Franco
Silvestri, Fabrizio
author_facet Bucarelli, Maria Sofia
D'Inverno, Giuseppe Alessio
Bianchini, Monica
Scarselli, Franco
Silvestri, Fabrizio
contents In the context of deep learning models, attention has recently been paid to studying the surface of the loss function in order to better understand training with methods based on gradient descent. This search for an appropriate description, both analytical and topological, has led to numerous efforts to identify spurious minima and characterize gradient dynamics. Our work aims to contribute to this field by providing a topological measure to evaluate loss complexity in the case of multilayer neural networks. We compare deep and shallow architectures with common sigmoidal activation functions by deriving upper and lower bounds on the complexity of their loss function and revealing how that complexity is influenced by the number of hidden units, training models, and the activation function used. Additionally, we found that certain variations in the loss function or model architecture, such as adding an $\ell_2$ regularization term or implementing skip connections in a feedforward network, do not affect loss topology in specific cases.
format Preprint
id arxiv_https___arxiv_org_abs_2401_03824
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A topological description of loss surfaces based on Betti Numbers
Bucarelli, Maria Sofia
D'Inverno, Giuseppe Alessio
Bianchini, Monica
Scarselli, Franco
Silvestri, Fabrizio
Machine Learning
In the context of deep learning models, attention has recently been paid to studying the surface of the loss function in order to better understand training with methods based on gradient descent. This search for an appropriate description, both analytical and topological, has led to numerous efforts to identify spurious minima and characterize gradient dynamics. Our work aims to contribute to this field by providing a topological measure to evaluate loss complexity in the case of multilayer neural networks. We compare deep and shallow architectures with common sigmoidal activation functions by deriving upper and lower bounds on the complexity of their loss function and revealing how that complexity is influenced by the number of hidden units, training models, and the activation function used. Additionally, we found that certain variations in the loss function or model architecture, such as adding an $\ell_2$ regularization term or implementing skip connections in a feedforward network, do not affect loss topology in specific cases.
title A topological description of loss surfaces based on Betti Numbers
topic Machine Learning
url https://arxiv.org/abs/2401.03824