A Hybrid Mixture Approach for Clustering and Characterizing Cancer Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kareem, Kazeem, Dai, Fan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909696271581184
author Kareem, Kazeem
Dai, Fan
author_facet Kareem, Kazeem
Dai, Fan
contents Model-based clustering is widely used for identifying and distinguishing types of diseases. However, modern biomedical data coming with high dimensions make it challenging to perform the model estimation in traditional cluster analysis. The incorporation of factor analyzer into the mixture model provides a way to characterize the large set of data features, but the current estimation method is computationally impractical for massive data due to the intrinsic slow convergence of the embedded algorithms, and the incapability to vary the size of the factor analyzers, preventing the implementation of a generalized mixture of factor analyzers and further characterization of the data clusters. We propose a hybrid matrix-free computational scheme to efficiently estimate the clusters and model parameters based on a Gaussian mixture along with generalized factor analyzers to summarize the large number of variables using a small set of underlying factors. Our approach outperforms the existing method with faster convergence while maintaining high clustering accuracy. Our algorithms are applied to accurately identify and distinguish types of breast cancer based on large tumor samples, and to provide a generalized characterization for subtypes of lymphoma using massive gene records.
format Preprint
id arxiv_https___arxiv_org_abs_2507_14380
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Hybrid Mixture Approach for Clustering and Characterizing Cancer Data
Kareem, Kazeem
Dai, Fan
Methodology
Tissues and Organs
Applications
Computation
Machine Learning
62H05, 62H12, 62H20, 62H25, 62H30, 62P10
G.3; I.2; I.5; I.6; J.3
Model-based clustering is widely used for identifying and distinguishing types of diseases. However, modern biomedical data coming with high dimensions make it challenging to perform the model estimation in traditional cluster analysis. The incorporation of factor analyzer into the mixture model provides a way to characterize the large set of data features, but the current estimation method is computationally impractical for massive data due to the intrinsic slow convergence of the embedded algorithms, and the incapability to vary the size of the factor analyzers, preventing the implementation of a generalized mixture of factor analyzers and further characterization of the data clusters. We propose a hybrid matrix-free computational scheme to efficiently estimate the clusters and model parameters based on a Gaussian mixture along with generalized factor analyzers to summarize the large number of variables using a small set of underlying factors. Our approach outperforms the existing method with faster convergence while maintaining high clustering accuracy. Our algorithms are applied to accurately identify and distinguish types of breast cancer based on large tumor samples, and to provide a generalized characterization for subtypes of lymphoma using massive gene records.
title A Hybrid Mixture Approach for Clustering and Characterizing Cancer Data
topic Methodology
Tissues and Organs
Applications
Computation
Machine Learning
62H05, 62H12, 62H20, 62H25, 62H30, 62P10
G.3; I.2; I.5; I.6; J.3
url https://arxiv.org/abs/2507.14380