ALPCAHUS: Subspace Clustering for Heteroscedastic Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cavazos, Javier Salazar, Fessler, Jeffrey A, Balzano, Laura
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917210619904000
author Cavazos, Javier Salazar
Fessler, Jeffrey A
Balzano, Laura
author_facet Cavazos, Javier Salazar
Fessler, Jeffrey A
Balzano, Laura
contents Principal component analysis (PCA) is a key tool in the field of data dimensionality reduction. Various methods have been proposed to extend PCA to the union of subspace (UoS) setting for clustering data that comes from multiple subspaces like K-Subspaces (KSS). However, some applications involve heterogeneous data that vary in quality due to noise characteristics associated with each data sample. Heteroscedastic methods aim to deal with such mixed data quality. This paper develops a heteroscedastic-based subspace clustering method, named ALPCAHUS, that can estimate the sample-wise noise variances and use this information to improve the estimate of the subspace bases associated with the low-rank structure of the data. This clustering algorithm builds on K-Subspaces (KSS) principles by extending the recently proposed heteroscedastic PCA method, named LR-ALPCAH, for clusters with heteroscedastic noise in the UoS setting. Simulations and real-data experiments show the effectiveness of accounting for data heteroscedasticity compared to existing clustering algorithms. Code available at https://github.com/javiersc1/ALPCAHUS.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18918
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ALPCAHUS: Subspace Clustering for Heteroscedastic Data
Cavazos, Javier Salazar
Fessler, Jeffrey A
Balzano, Laura
Machine Learning
Signal Processing
Principal component analysis (PCA) is a key tool in the field of data dimensionality reduction. Various methods have been proposed to extend PCA to the union of subspace (UoS) setting for clustering data that comes from multiple subspaces like K-Subspaces (KSS). However, some applications involve heterogeneous data that vary in quality due to noise characteristics associated with each data sample. Heteroscedastic methods aim to deal with such mixed data quality. This paper develops a heteroscedastic-based subspace clustering method, named ALPCAHUS, that can estimate the sample-wise noise variances and use this information to improve the estimate of the subspace bases associated with the low-rank structure of the data. This clustering algorithm builds on K-Subspaces (KSS) principles by extending the recently proposed heteroscedastic PCA method, named LR-ALPCAH, for clusters with heteroscedastic noise in the UoS setting. Simulations and real-data experiments show the effectiveness of accounting for data heteroscedasticity compared to existing clustering algorithms. Code available at https://github.com/javiersc1/ALPCAHUS.
title ALPCAHUS: Subspace Clustering for Heteroscedastic Data
topic Machine Learning
Signal Processing
url https://arxiv.org/abs/2505.18918