A Distributed Approach for Persistent Homology Computation on a Large Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ceccaroni, Riccardo, Di Rocco, Lorenzo, Petrillo, Umberto Ferraro, Brutti, Pierpaolo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914750437261312
author Ceccaroni, Riccardo
Di Rocco, Lorenzo
Petrillo, Umberto Ferraro
Brutti, Pierpaolo
author_facet Ceccaroni, Riccardo
Di Rocco, Lorenzo
Petrillo, Umberto Ferraro
Brutti, Pierpaolo
contents Persistent homology (PH) is a powerful mathematical method to automatically extract relevant insights from images, such as those obtained by high-resolution imaging devices like electron microscopes or new-generation telescopes. However, the application of this method comes at a very high computational cost, that is bound to explode more because new imaging devices generate an ever-growing amount of data. In this paper we present PixHomology, a novel algorithm for efficiently computing $0$-dimensional PH on 2D images, optimizing memory and processing time. By leveraging the Apache Spark framework, we also present a distributed version of our algorithm with several optimized variants, able to concurrently process large batches of astronomical images. Finally, we present the results of an experimental analysis showing that our algorithm and its distributed version are efficient in terms of required memory, execution time, and scalability, consistently outperforming existing state-of-the-art PH computation tools when used to process large datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2404_08245
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Distributed Approach for Persistent Homology Computation on a Large Scale
Ceccaroni, Riccardo
Di Rocco, Lorenzo
Petrillo, Umberto Ferraro
Brutti, Pierpaolo
Distributed, Parallel, and Cluster Computing
Computation
Persistent homology (PH) is a powerful mathematical method to automatically extract relevant insights from images, such as those obtained by high-resolution imaging devices like electron microscopes or new-generation telescopes. However, the application of this method comes at a very high computational cost, that is bound to explode more because new imaging devices generate an ever-growing amount of data. In this paper we present PixHomology, a novel algorithm for efficiently computing $0$-dimensional PH on 2D images, optimizing memory and processing time. By leveraging the Apache Spark framework, we also present a distributed version of our algorithm with several optimized variants, able to concurrently process large batches of astronomical images. Finally, we present the results of an experimental analysis showing that our algorithm and its distributed version are efficient in terms of required memory, execution time, and scalability, consistently outperforming existing state-of-the-art PH computation tools when used to process large datasets.
title A Distributed Approach for Persistent Homology Computation on a Large Scale
topic Distributed, Parallel, and Cluster Computing
Computation
url https://arxiv.org/abs/2404.08245