Loss Barcode: A Topological Measure of Escapability in Loss Landscapes
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2020
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910039361454080 |
|---|---|
| author | Barannikov, Serguei Voronkova, Daria Mironenko, Alexander Trofimov, Ilya Korotin, Alexander Sotnikov, Grigorii Burnaev, Evgeny |
| author_facet | Barannikov, Serguei Voronkova, Daria Mironenko, Alexander Trofimov, Ilya Korotin, Alexander Sotnikov, Grigorii Burnaev, Evgeny |
| contents | Neural network training is commonly based on SGD. However, the understanding of SGD's ability to converge to good local minima, given the non-convex nature of loss functions and the intricate geometric characteristics of loss landscapes, remains limited. In this paper, we apply topological data analysis methods to loss landscapes to gain insights into the learning process and generalization properties of deep neural networks. We use the loss function topology to relate the local behavior of gradient descent trajectories with the global properties of the loss surface. For this purpose, we define the neural network's Topological Obstructions score ("TO-score") with the help of robust topological invariants, barcodes of the loss function, which quantify the escapability of local minima for gradient-based optimization. Our two principal observations are: 1) the loss barcode of the neural network decreases with increasing depth and width, therefore the topological obstructions to learning diminish; 2) in certain situations there is a connection between the length of minima segments in the loss barcode and the minima's generalization errors. Our statements are based on extensive experiments with fully connected, convolutional, and transformer architectures and several datasets including MNIST, FMNIST, CIFAR10, CIFAR100, SVHN, and multilingual OSCAR text dataset. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2012_15834 |
| institution | arXiv |
| publishDate | 2020 |
| record_format | arxiv |
| spellingShingle | Loss Barcode: A Topological Measure of Escapability in Loss Landscapes Barannikov, Serguei Voronkova, Daria Mironenko, Alexander Trofimov, Ilya Korotin, Alexander Sotnikov, Grigorii Burnaev, Evgeny Machine Learning Artificial Intelligence Information Theory Dynamical Systems Neural network training is commonly based on SGD. However, the understanding of SGD's ability to converge to good local minima, given the non-convex nature of loss functions and the intricate geometric characteristics of loss landscapes, remains limited. In this paper, we apply topological data analysis methods to loss landscapes to gain insights into the learning process and generalization properties of deep neural networks. We use the loss function topology to relate the local behavior of gradient descent trajectories with the global properties of the loss surface. For this purpose, we define the neural network's Topological Obstructions score ("TO-score") with the help of robust topological invariants, barcodes of the loss function, which quantify the escapability of local minima for gradient-based optimization. Our two principal observations are: 1) the loss barcode of the neural network decreases with increasing depth and width, therefore the topological obstructions to learning diminish; 2) in certain situations there is a connection between the length of minima segments in the loss barcode and the minima's generalization errors. Our statements are based on extensive experiments with fully connected, convolutional, and transformer architectures and several datasets including MNIST, FMNIST, CIFAR10, CIFAR100, SVHN, and multilingual OSCAR text dataset. |
| title | Loss Barcode: A Topological Measure of Escapability in Loss Landscapes |
| topic | Machine Learning Artificial Intelligence Information Theory Dynamical Systems |
| url | https://arxiv.org/abs/2012.15834 |