Scalable Bayesian Clustering for Integrative Analysis of Multi-View Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cabral, Rafael, de Iorio, Maria, Harris, Andrew
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910584345198592
author Cabral, Rafael
de Iorio, Maria
Harris, Andrew
author_facet Cabral, Rafael
de Iorio, Maria
Harris, Andrew
contents In the era of Big Data, scalable and accurate clustering algorithms for high-dimensional data are essential. We present new Bayesian Distance Clustering (BDC) models and inference algorithms with improved scalability while maintaining the predictive accuracy of modern Bayesian non-parametric models. Unlike traditional methods, BDC models the distance between observations rather than the observations directly, offering a compromise between the scalability of distance-based methods and the enhanced predictive power and probabilistic interpretation of model-based methods. However, existing BDC models still rely on performing inference on the partition model to group observations into clusters. The support of this partition model grows exponentially with the dataset's size, complicating posterior space exploration and leading to many costly likelihood evaluations. Inspired by K-medoids, we propose using tessellations in discrete space to simplify inference by focusing the learning task on finding the best tessellation centers, or "medoids." Additionally, we extend our models to effectively handle multi-view data, such as data comprised of clusters that evolve across time, enhancing their applicability to complex datasets. The real data application in numismatics demonstrates the efficacy of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2408_17153
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scalable Bayesian Clustering for Integrative Analysis of Multi-View Data
Cabral, Rafael
de Iorio, Maria
Harris, Andrew
Methodology
Computation
62F15, 62H30, 68T05, 62P25
G.3; I.5.3; I.5.1; I.2.6
In the era of Big Data, scalable and accurate clustering algorithms for high-dimensional data are essential. We present new Bayesian Distance Clustering (BDC) models and inference algorithms with improved scalability while maintaining the predictive accuracy of modern Bayesian non-parametric models. Unlike traditional methods, BDC models the distance between observations rather than the observations directly, offering a compromise between the scalability of distance-based methods and the enhanced predictive power and probabilistic interpretation of model-based methods. However, existing BDC models still rely on performing inference on the partition model to group observations into clusters. The support of this partition model grows exponentially with the dataset's size, complicating posterior space exploration and leading to many costly likelihood evaluations. Inspired by K-medoids, we propose using tessellations in discrete space to simplify inference by focusing the learning task on finding the best tessellation centers, or "medoids." Additionally, we extend our models to effectively handle multi-view data, such as data comprised of clusters that evolve across time, enhancing their applicability to complex datasets. The real data application in numismatics demonstrates the efficacy of our approach.
title Scalable Bayesian Clustering for Integrative Analysis of Multi-View Data
topic Methodology
Computation
62F15, 62H30, 68T05, 62P25
G.3; I.5.3; I.5.1; I.2.6
url https://arxiv.org/abs/2408.17153