Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge Ensembles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ruben, Benjamin S., Pehlevan, Cengiz
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914636338561024
author Ruben, Benjamin S.
Pehlevan, Cengiz
author_facet Ruben, Benjamin S.
Pehlevan, Cengiz
contents Feature bagging is a well-established ensembling method which aims to reduce prediction variance by combining predictions of many estimators trained on subsets or projections of features. Here, we develop a theory of feature-bagging in noisy least-squares ridge ensembles and simplify the resulting learning curves in the special case of equicorrelated data. Using analytical learning curves, we demonstrate that subsampling shifts the double-descent peak of a linear predictor. This leads us to introduce heterogeneous feature ensembling, with estimators built on varying numbers of feature dimensions, as a computationally efficient method to mitigate double-descent. Then, we compare the performance of a feature-subsampling ensemble to a single linear predictor, describing a trade-off between noise amplification due to subsampling and noise reduction due to ensembling. Our qualitative insights carry over to linear classifiers applied to image classification tasks with realistic datasets constructed using a state-of-the-art deep learning feature map.
format Preprint
id arxiv_https___arxiv_org_abs_2307_03176
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge Ensembles
Ruben, Benjamin S.
Pehlevan, Cengiz
Machine Learning
Disordered Systems and Neural Networks
Neurons and Cognition
Feature bagging is a well-established ensembling method which aims to reduce prediction variance by combining predictions of many estimators trained on subsets or projections of features. Here, we develop a theory of feature-bagging in noisy least-squares ridge ensembles and simplify the resulting learning curves in the special case of equicorrelated data. Using analytical learning curves, we demonstrate that subsampling shifts the double-descent peak of a linear predictor. This leads us to introduce heterogeneous feature ensembling, with estimators built on varying numbers of feature dimensions, as a computationally efficient method to mitigate double-descent. Then, we compare the performance of a feature-subsampling ensemble to a single linear predictor, describing a trade-off between noise amplification due to subsampling and noise reduction due to ensembling. Our qualitative insights carry over to linear classifiers applied to image classification tasks with realistic datasets constructed using a state-of-the-art deep learning feature map.
title Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge Ensembles
topic Machine Learning
Disordered Systems and Neural Networks
Neurons and Cognition
url https://arxiv.org/abs/2307.03176