An Asynchronous Distributed-Memory Parallel Algorithm for k-mer Counting

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hati, Souvadra, Hayashi, Akihiro, Vuduc, Richard
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909603470508032
author Hati, Souvadra
Hayashi, Akihiro
Vuduc, Richard
author_facet Hati, Souvadra
Hayashi, Akihiro
Vuduc, Richard
contents This paper describes a new asynchronous algorithm and implementation for the problem of k-mer counting (KC), which concerns quantifying the frequency of length k substrings in a DNA sequence. This operation is common to many computational biology workloads and can take up to 77% of the total runtime of de novo genome assembly. The performance and scalability of the current state-of-the-art distributed-memory KC algorithm are hampered by multiple rounds of Many-To-Many collectives. Therefore, we develop an asynchronous algorithm (DAKC) that uses fine-grained, asynchronous messages to obviate most of this global communication while utilizing network bandwidth efficiently via custom message aggregation protocols. DAKC can perform strong scaling up to 256 nodes (512 sockets / 6K cores) and can count k-mers up to 9x faster than the state-of-the-art distributed-memory algorithm, and up to 100x faster than the shared-memory alternative. We also provide an analytical model to understand the hardware resource utilization of our asynchronous KC algorithm and provide insights on the performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04431
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Asynchronous Distributed-Memory Parallel Algorithm for k-mer Counting
Hati, Souvadra
Hayashi, Akihiro
Vuduc, Richard
Distributed, Parallel, and Cluster Computing
Genomics
This paper describes a new asynchronous algorithm and implementation for the problem of k-mer counting (KC), which concerns quantifying the frequency of length k substrings in a DNA sequence. This operation is common to many computational biology workloads and can take up to 77% of the total runtime of de novo genome assembly. The performance and scalability of the current state-of-the-art distributed-memory KC algorithm are hampered by multiple rounds of Many-To-Many collectives. Therefore, we develop an asynchronous algorithm (DAKC) that uses fine-grained, asynchronous messages to obviate most of this global communication while utilizing network bandwidth efficiently via custom message aggregation protocols. DAKC can perform strong scaling up to 256 nodes (512 sockets / 6K cores) and can count k-mers up to 9x faster than the state-of-the-art distributed-memory algorithm, and up to 100x faster than the shared-memory alternative. We also provide an analytical model to understand the hardware resource utilization of our asynchronous KC algorithm and provide insights on the performance.
title An Asynchronous Distributed-Memory Parallel Algorithm for k-mer Counting
topic Distributed, Parallel, and Cluster Computing
Genomics
url https://arxiv.org/abs/2505.04431