Scalable Co-Clustering for Large-Scale Data through Dynamic Partitioning and Hierarchical Merging

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Zihan, Huang, Zhaoke, Yan, Hong
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909541741887488
author Wu, Zihan
Huang, Zhaoke
Yan, Hong
author_facet Wu, Zihan
Huang, Zhaoke
Yan, Hong
contents Co-clustering simultaneously clusters rows and columns, revealing more fine-grained groups. However, existing co-clustering methods suffer from poor scalability and cannot handle large-scale data. This paper presents a novel and scalable co-clustering method designed to uncover intricate patterns in high-dimensional, large-scale datasets. Specifically, we first propose a large matrix partitioning algorithm that partitions a large matrix into smaller submatrices, enabling parallel co-clustering. This method employs a probabilistic model to optimize the configuration of submatrices, balancing the computational efficiency and depth of analysis. Additionally, we propose a hierarchical co-cluster merging algorithm that efficiently identifies and merges co-clusters from these submatrices, enhancing the robustness and reliability of the process. Extensive evaluations validate the effectiveness and efficiency of our method. Experimental results demonstrate a significant reduction in computation time, with an approximate 83% decrease for dense matrices and up to 30% for sparse matrices.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scalable Co-Clustering for Large-Scale Data through Dynamic Partitioning and Hierarchical Merging
Wu, Zihan
Huang, Zhaoke
Yan, Hong
Distributed, Parallel, and Cluster Computing
Machine Learning
H.2.8
Co-clustering simultaneously clusters rows and columns, revealing more fine-grained groups. However, existing co-clustering methods suffer from poor scalability and cannot handle large-scale data. This paper presents a novel and scalable co-clustering method designed to uncover intricate patterns in high-dimensional, large-scale datasets. Specifically, we first propose a large matrix partitioning algorithm that partitions a large matrix into smaller submatrices, enabling parallel co-clustering. This method employs a probabilistic model to optimize the configuration of submatrices, balancing the computational efficiency and depth of analysis. Additionally, we propose a hierarchical co-cluster merging algorithm that efficiently identifies and merges co-clusters from these submatrices, enhancing the robustness and reliability of the process. Extensive evaluations validate the effectiveness and efficiency of our method. Experimental results demonstrate a significant reduction in computation time, with an approximate 83% decrease for dense matrices and up to 30% for sparse matrices.
title Scalable Co-Clustering for Large-Scale Data through Dynamic Partitioning and Hierarchical Merging
topic Distributed, Parallel, and Cluster Computing
Machine Learning
H.2.8
url https://arxiv.org/abs/2410.18113