Saved in:
Bibliographic Details
Main Authors: Martinetz, Julius, Linse, Christoph, Martinetz, Thomas
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.16868
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929554164023296
author Martinetz, Julius
Linse, Christoph
Martinetz, Thomas
author_facet Martinetz, Julius
Linse, Christoph
Martinetz, Thomas
contents We investigate the learning dynamics of classifiers in scenarios where classes are separable or classifiers are over-parameterized. In both cases, Empirical Risk Minimization (ERM) results in zero training error. However, there are many global minima with a training error of zero, some of which generalize well and some of which do not. We show that in separable classes scenarios the proportion of "bad" global minima diminishes exponentially with the number of training data n. Our analysis provides bounds and learning curves dependent solely on the density distribution of the true error for the given classifier function set, irrespective of the set's size or complexity (e.g., number of parameters). This observation may shed light on the unexpectedly good generalization of over-parameterized Neural Networks. For the over-parameterized scenario, we propose a model for the density distribution of the true error, yielding learning curves that align with experiments on MNIST and CIFAR-10.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16868
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking generalization of classifiers in separable classes scenarios and over-parameterized regimes
Martinetz, Julius
Linse, Christoph
Martinetz, Thomas
Machine Learning
Computer Vision and Pattern Recognition
We investigate the learning dynamics of classifiers in scenarios where classes are separable or classifiers are over-parameterized. In both cases, Empirical Risk Minimization (ERM) results in zero training error. However, there are many global minima with a training error of zero, some of which generalize well and some of which do not. We show that in separable classes scenarios the proportion of "bad" global minima diminishes exponentially with the number of training data n. Our analysis provides bounds and learning curves dependent solely on the density distribution of the true error for the given classifier function set, irrespective of the set's size or complexity (e.g., number of parameters). This observation may shed light on the unexpectedly good generalization of over-parameterized Neural Networks. For the over-parameterized scenario, we propose a model for the density distribution of the true error, yielding learning curves that align with experiments on MNIST and CIFAR-10.
title Rethinking generalization of classifiers in separable classes scenarios and over-parameterized regimes
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.16868