A Unified Framework for Variable Selection in Model-Based Clustering with Missing Not at Random

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ho, Binh H., Chi, Long Nguyen, Nguyen, TrungTin, Nguyen, Binh T., Hoang, Van Ha, Drovandi, Christopher
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912686244102144
author Ho, Binh H.
Chi, Long Nguyen
Nguyen, TrungTin
Nguyen, Binh T.
Hoang, Van Ha
Drovandi, Christopher
author_facet Ho, Binh H.
Chi, Long Nguyen
Nguyen, TrungTin
Nguyen, Binh T.
Hoang, Van Ha
Drovandi, Christopher
contents Model-based clustering integrated with variable selection is a powerful tool for uncovering latent structures within complex data. However, its effectiveness is often hindered by challenges such as identifying relevant variables that define heterogeneous subgroups and handling data that are missing not at random, a prevalent issue in fields like transcriptomics. While several notable methods have been proposed to address these problems, they typically tackle each issue in isolation, thereby limiting their flexibility and adaptability. This paper introduces a unified framework designed to address these challenges simultaneously. Our approach incorporates a data-driven penalty matrix into penalized clustering to enable more flexible variable selection, along with a mechanism that explicitly models the relationship between missingness and latent class membership. We demonstrate that, under certain regularity conditions, the proposed framework achieves both asymptotic consistency and selection consistency, even in the presence of missing data. This unified strategy significantly enhances the capability and efficiency of model-based clustering, advancing methodologies for identifying informative variables that define homogeneous subgroups in the presence of complex missing data patterns. The performance of the framework, including its computational efficiency, is evaluated through simulations and demonstrated using both synthetic and real-world transcriptomic datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19093
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Unified Framework for Variable Selection in Model-Based Clustering with Missing Not at Random
Ho, Binh H.
Chi, Long Nguyen
Nguyen, TrungTin
Nguyen, Binh T.
Hoang, Van Ha
Drovandi, Christopher
Methodology
Machine Learning
Statistics Theory
Applications
Model-based clustering integrated with variable selection is a powerful tool for uncovering latent structures within complex data. However, its effectiveness is often hindered by challenges such as identifying relevant variables that define heterogeneous subgroups and handling data that are missing not at random, a prevalent issue in fields like transcriptomics. While several notable methods have been proposed to address these problems, they typically tackle each issue in isolation, thereby limiting their flexibility and adaptability. This paper introduces a unified framework designed to address these challenges simultaneously. Our approach incorporates a data-driven penalty matrix into penalized clustering to enable more flexible variable selection, along with a mechanism that explicitly models the relationship between missingness and latent class membership. We demonstrate that, under certain regularity conditions, the proposed framework achieves both asymptotic consistency and selection consistency, even in the presence of missing data. This unified strategy significantly enhances the capability and efficiency of model-based clustering, advancing methodologies for identifying informative variables that define homogeneous subgroups in the presence of complex missing data patterns. The performance of the framework, including its computational efficiency, is evaluated through simulations and demonstrated using both synthetic and real-world transcriptomic datasets.
title A Unified Framework for Variable Selection in Model-Based Clustering with Missing Not at Random
topic Methodology
Machine Learning
Statistics Theory
Applications
url https://arxiv.org/abs/2505.19093