LinearAlifold: Linear-Time Consensus Structure Prediction for RNA Alignments

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Malik, Apoorv, Zhang, Liang, Gautam, Milan, Dai, Ning, Li, Sizhen, Zhang, He, Mathews, David H., Huang, Liang
Format: Preprint
Publié: 2022
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916312359370752
author Malik, Apoorv
Zhang, Liang
Gautam, Milan
Dai, Ning
Li, Sizhen
Zhang, He
Mathews, David H.
Huang, Liang
author_facet Malik, Apoorv
Zhang, Liang
Gautam, Milan
Dai, Ning
Li, Sizhen
Zhang, He
Mathews, David H.
Huang, Liang
contents Predicting the consensus structure of a set of aligned RNA homologs is a convenient method to find conserved structures in an RNA genome, which has many applications including viral diagnostics and therapeutics. However, the most commonly used tool for this task, RNAalifold, is prohibitively slow for long sequences, due to a cubic scaling with the sequence length, taking over a day on 400 SARS-CoV-2 and SARS-related genomes (~30,000nt). We present LinearAlifold, a much faster alternative that scales linearly with both the sequence length and the number of sequences, based on our work LinearFold that folds a single RNA in linear time. Our work is orders of magnitude faster than RNAalifold (0.7 hours on the above 400 genomes, or ~36$\times$ speedup) and achieves higher accuracies when compared to a database of known structures. More interestingly, LinearAlifold's prediction on SARS-CoV-2 correlates well with experimentally determined structures, substantially outperforming RNAalifold. Finally, LinearAlifold supports two energy models (Vienna and BL*) and four modes: minimum free energy (MFE), maximum expected accuracy (MEA), ThreshKnot, and stochastic sampling, each of which takes under an hour for hundreds of SARS-CoV variants. Our resource is at: https://github.com/LinearFold/LinearAlifold (code) and http://linearfold.org/linear-alifold (server).
format Preprint
id arxiv_https___arxiv_org_abs_2206_14794
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle LinearAlifold: Linear-Time Consensus Structure Prediction for RNA Alignments
Malik, Apoorv
Zhang, Liang
Gautam, Milan
Dai, Ning
Li, Sizhen
Zhang, He
Mathews, David H.
Huang, Liang
Biomolecules
Data Structures and Algorithms
Biological Physics
Quantitative Methods
Predicting the consensus structure of a set of aligned RNA homologs is a convenient method to find conserved structures in an RNA genome, which has many applications including viral diagnostics and therapeutics. However, the most commonly used tool for this task, RNAalifold, is prohibitively slow for long sequences, due to a cubic scaling with the sequence length, taking over a day on 400 SARS-CoV-2 and SARS-related genomes (~30,000nt). We present LinearAlifold, a much faster alternative that scales linearly with both the sequence length and the number of sequences, based on our work LinearFold that folds a single RNA in linear time. Our work is orders of magnitude faster than RNAalifold (0.7 hours on the above 400 genomes, or ~36$\times$ speedup) and achieves higher accuracies when compared to a database of known structures. More interestingly, LinearAlifold's prediction on SARS-CoV-2 correlates well with experimentally determined structures, substantially outperforming RNAalifold. Finally, LinearAlifold supports two energy models (Vienna and BL*) and four modes: minimum free energy (MFE), maximum expected accuracy (MEA), ThreshKnot, and stochastic sampling, each of which takes under an hour for hundreds of SARS-CoV variants. Our resource is at: https://github.com/LinearFold/LinearAlifold (code) and http://linearfold.org/linear-alifold (server).
title LinearAlifold: Linear-Time Consensus Structure Prediction for RNA Alignments
topic Biomolecules
Data Structures and Algorithms
Biological Physics
Quantitative Methods
url https://arxiv.org/abs/2206.14794