Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Daiwei, Chang, Wei-Kai, Chaudhari, Pratik
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2305.17332
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914979238641664
author Chen, Daiwei
Chang, Wei-Kai
Chaudhari, Pratik
author_facet Chen, Daiwei
Chang, Wei-Kai
Chaudhari, Pratik
contents We use a formal correspondence between thermodynamics and inference, where the number of samples can be thought of as the inverse temperature, to study a quantity called ``learning capacity'' which is a measure of the effective dimensionality of a model. We show that the learning capacity is a useful notion of the complexity because (a) it correlates well with the test loss and it is a tiny fraction of the number of parameters for many deep networks trained on typical datasets, (b) it depends upon the number of samples used for training, (c) it is numerically consistent with notions of capacity obtained from PAC-Bayes generalization bounds, and (d) the test loss as a function of the learning capacity does not exhibit double descent. We show that the learning capacity saturates at very small and very large sample sizes; the threshold that characterizes the transition between these two regimes provides guidelines as to when one should procure more data and when one should search for a different architecture to improve performance. We show how the learning capacity can be used to provide a quantitative notion of capacity even for non-parametric models such as random forests and nearest neighbor classifiers.
format Preprint
id arxiv_https___arxiv_org_abs_2305_17332
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning Capacity: A Measure of the Effective Dimensionality of a Model
Chen, Daiwei
Chang, Wei-Kai
Chaudhari, Pratik
Machine Learning
Information Theory
We use a formal correspondence between thermodynamics and inference, where the number of samples can be thought of as the inverse temperature, to study a quantity called ``learning capacity'' which is a measure of the effective dimensionality of a model. We show that the learning capacity is a useful notion of the complexity because (a) it correlates well with the test loss and it is a tiny fraction of the number of parameters for many deep networks trained on typical datasets, (b) it depends upon the number of samples used for training, (c) it is numerically consistent with notions of capacity obtained from PAC-Bayes generalization bounds, and (d) the test loss as a function of the learning capacity does not exhibit double descent. We show that the learning capacity saturates at very small and very large sample sizes; the threshold that characterizes the transition between these two regimes provides guidelines as to when one should procure more data and when one should search for a different architecture to improve performance. We show how the learning capacity can be used to provide a quantitative notion of capacity even for non-parametric models such as random forests and nearest neighbor classifiers.
title Learning Capacity: A Measure of the Effective Dimensionality of a Model
topic Machine Learning
Information Theory
url https://arxiv.org/abs/2305.17332