On the (In)Significance of Feature Selection in High-Dimensional Datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Neekhra, Bhavesh, Gupta, Debayan, Chakrabarti, Partha Pratim
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916957875339264
author Neekhra, Bhavesh
Gupta, Debayan
Chakrabarti, Partha Pratim
author_facet Neekhra, Bhavesh
Gupta, Debayan
Chakrabarti, Partha Pratim
contents Feature selection (FS) is assumed to improve predictive performance and identify meaningful features in high-dimensional datasets. Surprisingly, small random subsets of features (0.02-1%) match or outperform the predictive performance of both full feature sets and FS across 28 out of 30 diverse datasets (microarray, bulk and single-cell RNA-Seq, mass spectrometry, imaging, etc.). In short, any arbitrary set of features is as good as any other (with surprisingly low variance in results) - so how can a particular set of selected features be "important" if they perform no better than an arbitrary set? These results challenge the assumption that computationally selected features reliably capture meaningful signals, emphasizing the importance of rigorous validation before interpreting selected features as actionable, particularly in computational genomics.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03593
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the (In)Significance of Feature Selection in High-Dimensional Datasets
Neekhra, Bhavesh
Gupta, Debayan
Chakrabarti, Partha Pratim
Machine Learning
Genomics
Feature selection (FS) is assumed to improve predictive performance and identify meaningful features in high-dimensional datasets. Surprisingly, small random subsets of features (0.02-1%) match or outperform the predictive performance of both full feature sets and FS across 28 out of 30 diverse datasets (microarray, bulk and single-cell RNA-Seq, mass spectrometry, imaging, etc.). In short, any arbitrary set of features is as good as any other (with surprisingly low variance in results) - so how can a particular set of selected features be "important" if they perform no better than an arbitrary set? These results challenge the assumption that computationally selected features reliably capture meaningful signals, emphasizing the importance of rigorous validation before interpreting selected features as actionable, particularly in computational genomics.
title On the (In)Significance of Feature Selection in High-Dimensional Datasets
topic Machine Learning
Genomics
url https://arxiv.org/abs/2508.03593