MirLibSpark: A Scalable NGS Plant MicroRNA Prediction Pipeline for Multi-Library Functional Annotation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Chao-Jung, Remita, Amine M., Diallo, Abdoulaye Baniré
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915129221709824
author Wu, Chao-Jung
Remita, Amine M.
Diallo, Abdoulaye Baniré
author_facet Wu, Chao-Jung
Remita, Amine M.
Diallo, Abdoulaye Baniré
contents The emergence of the Next Generation Sequencing increases drastically the volume of transcriptomic data. Although many standalone algorithms and workflows for novel microRNA (miRNA) prediction have been proposed, few are designed for processing large volume of sequence data from large genomes, and even fewer further annotate functional miRNAs by analyzing multiple libraries. We propose an improved pipeline for a high volume data facility by implementing mirLibSpark based on the Apache Spark framework. This pipeline is the fastest actual method, and provides an accuracy improvement compared to the standard. In this paper, we deliver the first distributed functional miRNA predictor as a standalone and fully automated package. It is an efficient and accurate miRNA predictor with functional insight. Furthermore, it compiles with the gold-standard requirement on plant miRNA predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17998
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MirLibSpark: A Scalable NGS Plant MicroRNA Prediction Pipeline for Multi-Library Functional Annotation
Wu, Chao-Jung
Remita, Amine M.
Diallo, Abdoulaye Baniré
Distributed, Parallel, and Cluster Computing
Genomics
J.3; I.5.3; D.2.11
The emergence of the Next Generation Sequencing increases drastically the volume of transcriptomic data. Although many standalone algorithms and workflows for novel microRNA (miRNA) prediction have been proposed, few are designed for processing large volume of sequence data from large genomes, and even fewer further annotate functional miRNAs by analyzing multiple libraries. We propose an improved pipeline for a high volume data facility by implementing mirLibSpark based on the Apache Spark framework. This pipeline is the fastest actual method, and provides an accuracy improvement compared to the standard. In this paper, we deliver the first distributed functional miRNA predictor as a standalone and fully automated package. It is an efficient and accurate miRNA predictor with functional insight. Furthermore, it compiles with the gold-standard requirement on plant miRNA predictions.
title MirLibSpark: A Scalable NGS Plant MicroRNA Prediction Pipeline for Multi-Library Functional Annotation
topic Distributed, Parallel, and Cluster Computing
Genomics
J.3; I.5.3; D.2.11
url https://arxiv.org/abs/2501.17998