A Distributed Block Chebyshev-Davidson Algorithm for Parallel Spectral Clustering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pang, Qiyuan, Yang, Haizhao
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911748427087872
author Pang, Qiyuan
Yang, Haizhao
author_facet Pang, Qiyuan
Yang, Haizhao
contents We develop a distributed Block Chebyshev-Davidson algorithm to solve large-scale leading eigenvalue problems for spectral analysis in spectral clustering. First, the efficiency of the Chebyshev-Davidson algorithm relies on the prior knowledge of the eigenvalue spectrum, which could be expensive to estimate. This issue can be lessened by the analytic spectrum estimation of the Laplacian or normalized Laplacian matrices in spectral clustering, making the proposed algorithm very efficient for spectral clustering. Second, to make the proposed algorithm capable of analyzing big data, a distributed and parallel version has been developed with attractive scalability. The speedup by parallel computing is approximately equivalent to $\sqrt{p}$, where $p$ denotes the number of processes. {Numerical results will be provided to demonstrate its efficiency in spectral clustering and scalability advantage over existing eigensolvers used for spectral clustering in parallel computing environments.}
format Preprint
id arxiv_https___arxiv_org_abs_2212_04443
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle A Distributed Block Chebyshev-Davidson Algorithm for Parallel Spectral Clustering
Pang, Qiyuan
Yang, Haizhao
Machine Learning
Distributed, Parallel, and Cluster Computing
Numerical Analysis
We develop a distributed Block Chebyshev-Davidson algorithm to solve large-scale leading eigenvalue problems for spectral analysis in spectral clustering. First, the efficiency of the Chebyshev-Davidson algorithm relies on the prior knowledge of the eigenvalue spectrum, which could be expensive to estimate. This issue can be lessened by the analytic spectrum estimation of the Laplacian or normalized Laplacian matrices in spectral clustering, making the proposed algorithm very efficient for spectral clustering. Second, to make the proposed algorithm capable of analyzing big data, a distributed and parallel version has been developed with attractive scalability. The speedup by parallel computing is approximately equivalent to $\sqrt{p}$, where $p$ denotes the number of processes. {Numerical results will be provided to demonstrate its efficiency in spectral clustering and scalability advantage over existing eigensolvers used for spectral clustering in parallel computing environments.}
title A Distributed Block Chebyshev-Davidson Algorithm for Parallel Spectral Clustering
topic Machine Learning
Distributed, Parallel, and Cluster Computing
Numerical Analysis
url https://arxiv.org/abs/2212.04443