Bayesian clustering of high-dimensional data via latent repulsive mixtures

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ghilotti, Lorenzo, Beraha, Mario, Guglielmi, Alessandra
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911897337462784
author Ghilotti, Lorenzo
Beraha, Mario
Guglielmi, Alessandra
author_facet Ghilotti, Lorenzo
Beraha, Mario
Guglielmi, Alessandra
contents Model-based clustering of moderate or large dimensional data is notoriously difficult. We propose a model for simultaneous dimensionality reduction and clustering by assuming a mixture model for a set of latent scores, which are then linked to the observations via a Gaussian latent factor model. This approach was recently investigated by Chandra et al. (2023). The authors use a factor-analytic representation and assume a mixture model for the latent factors. However, performance can deteriorate in the presence of model misspecification. Assuming a repulsive point process prior for the component-specific means of the mixture for the latent scores is shown to yield a more robust model that outperforms the standard mixture model for the latent factors in several simulated scenarios. The repulsive point process must be anisotropic to favor well-separated clusters of data, and its density should be tractable for efficient posterior inference. We address these issues by proposing a general construction for anisotropic determinantal point processes. We illustrate our model in simulations as well as a plant species co-occurrence dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2303_02438
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Bayesian clustering of high-dimensional data via latent repulsive mixtures
Ghilotti, Lorenzo
Beraha, Mario
Guglielmi, Alessandra
Methodology
Model-based clustering of moderate or large dimensional data is notoriously difficult. We propose a model for simultaneous dimensionality reduction and clustering by assuming a mixture model for a set of latent scores, which are then linked to the observations via a Gaussian latent factor model. This approach was recently investigated by Chandra et al. (2023). The authors use a factor-analytic representation and assume a mixture model for the latent factors. However, performance can deteriorate in the presence of model misspecification. Assuming a repulsive point process prior for the component-specific means of the mixture for the latent scores is shown to yield a more robust model that outperforms the standard mixture model for the latent factors in several simulated scenarios. The repulsive point process must be anisotropic to favor well-separated clusters of data, and its density should be tractable for efficient posterior inference. We address these issues by proposing a general construction for anisotropic determinantal point processes. We illustrate our model in simulations as well as a plant species co-occurrence dataset.
title Bayesian clustering of high-dimensional data via latent repulsive mixtures
topic Methodology
url https://arxiv.org/abs/2303.02438