FINDER: Feature Inference on Noisy Datasets using Eigenspace Residuals

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Murphy, Trajan, Dogra, Akshunna S., Gu, Hanfeng, Meredith, Caleb, Kon, Mark, Castrillion-Candas, Julio Enrique
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911227708440576
author Murphy, Trajan
Dogra, Akshunna S.
Gu, Hanfeng
Meredith, Caleb
Kon, Mark
Castrillion-Candas, Julio Enrique
author_facet Murphy, Trajan
Dogra, Akshunna S.
Gu, Hanfeng
Meredith, Caleb
Kon, Mark
Castrillion-Candas, Julio Enrique
contents ''Noisy'' datasets (regimes with low signal to noise ratios, small sample sizes, faulty data collection, etc) remain a key research frontier for classification methods with both theoretical and practical implications. We introduce FINDER, a rigorous framework for analyzing generic classification problems, with tailored algorithms for noisy datasets. FINDER incorporates fundamental stochastic analysis ideas into the feature learning and inference stages to optimally account for the randomness inherent to all empirical datasets. We construct ''stochastic features'' by first viewing empirical datasets as realizations from an underlying random field (without assumptions on its exact distribution) and then mapping them to appropriate Hilbert spaces. The Kosambi-Karhunen-Loéve expansion (KLE) breaks these stochastic features into computable irreducible components, which allow classification over noisy datasets via an eigen-decomposition: data from different classes resides in distinct regions, identified by analyzing the spectrum of the associated operators. We validate FINDER on several challenging, data-deficient scientific domains, producing state of the art breakthroughs in: (i) Alzheimer's Disease stage classification, (ii) Remote sensing detection of deforestation. We end with a discussion on when FINDER is expected to outperform existing methods, its failure modes, and other limitations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19917
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FINDER: Feature Inference on Noisy Datasets using Eigenspace Residuals
Murphy, Trajan
Dogra, Akshunna S.
Gu, Hanfeng
Meredith, Caleb
Kon, Mark
Castrillion-Candas, Julio Enrique
Machine Learning
Computer Vision and Pattern Recognition
Numerical Analysis
''Noisy'' datasets (regimes with low signal to noise ratios, small sample sizes, faulty data collection, etc) remain a key research frontier for classification methods with both theoretical and practical implications. We introduce FINDER, a rigorous framework for analyzing generic classification problems, with tailored algorithms for noisy datasets. FINDER incorporates fundamental stochastic analysis ideas into the feature learning and inference stages to optimally account for the randomness inherent to all empirical datasets. We construct ''stochastic features'' by first viewing empirical datasets as realizations from an underlying random field (without assumptions on its exact distribution) and then mapping them to appropriate Hilbert spaces. The Kosambi-Karhunen-Loéve expansion (KLE) breaks these stochastic features into computable irreducible components, which allow classification over noisy datasets via an eigen-decomposition: data from different classes resides in distinct regions, identified by analyzing the spectrum of the associated operators. We validate FINDER on several challenging, data-deficient scientific domains, producing state of the art breakthroughs in: (i) Alzheimer's Disease stage classification, (ii) Remote sensing detection of deforestation. We end with a discussion on when FINDER is expected to outperform existing methods, its failure modes, and other limitations.
title FINDER: Feature Inference on Noisy Datasets using Eigenspace Residuals
topic Machine Learning
Computer Vision and Pattern Recognition
Numerical Analysis
url https://arxiv.org/abs/2510.19917