Deciphering RNA Secondary Structure Prediction: A Probabilistic K-Rook Matching Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Cheng, Gao, Zhangyang, Cao, Hanqun, Chen, Xingran, Wang, Ge, Wu, Lirong, Xia, Jun, Zheng, Jiangbin, Li, Stan Z.
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911925299838976
author Tan, Cheng
Gao, Zhangyang
Cao, Hanqun
Chen, Xingran
Wang, Ge
Wu, Lirong
Xia, Jun
Zheng, Jiangbin
Li, Stan Z.
author_facet Tan, Cheng
Gao, Zhangyang
Cao, Hanqun
Chen, Xingran
Wang, Ge
Wu, Lirong
Xia, Jun
Zheng, Jiangbin
Li, Stan Z.
contents The secondary structure of ribonucleic acid (RNA) is more stable and accessible in the cell than its tertiary structure, making it essential for functional prediction. Although deep learning has shown promising results in this field, current methods suffer from poor generalization and high complexity. In this work, we reformulate the RNA secondary structure prediction as a K-Rook problem, thereby simplifying the prediction process into probabilistic matching within a finite solution space. Building on this innovative perspective, we introduce RFold, a simple yet effective method that learns to predict the most matching K-Rook solution from the given sequence. RFold employs a bi-dimensional optimization strategy that decomposes the probabilistic matching problem into row-wise and column-wise components to reduce the matching complexity, simplifying the solving process while guaranteeing the validity of the output. Extensive experiments demonstrate that RFold achieves competitive performance and about eight times faster inference efficiency than the state-of-the-art approaches. The code and Colab demo are available in (http://github.com/A4Bio/RFold).
format Preprint
id arxiv_https___arxiv_org_abs_2212_14041
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Deciphering RNA Secondary Structure Prediction: A Probabilistic K-Rook Matching Perspective
Tan, Cheng
Gao, Zhangyang
Cao, Hanqun
Chen, Xingran
Wang, Ge
Wu, Lirong
Xia, Jun
Zheng, Jiangbin
Li, Stan Z.
Biomolecules
Artificial Intelligence
Machine Learning
The secondary structure of ribonucleic acid (RNA) is more stable and accessible in the cell than its tertiary structure, making it essential for functional prediction. Although deep learning has shown promising results in this field, current methods suffer from poor generalization and high complexity. In this work, we reformulate the RNA secondary structure prediction as a K-Rook problem, thereby simplifying the prediction process into probabilistic matching within a finite solution space. Building on this innovative perspective, we introduce RFold, a simple yet effective method that learns to predict the most matching K-Rook solution from the given sequence. RFold employs a bi-dimensional optimization strategy that decomposes the probabilistic matching problem into row-wise and column-wise components to reduce the matching complexity, simplifying the solving process while guaranteeing the validity of the output. Extensive experiments demonstrate that RFold achieves competitive performance and about eight times faster inference efficiency than the state-of-the-art approaches. The code and Colab demo are available in (http://github.com/A4Bio/RFold).
title Deciphering RNA Secondary Structure Prediction: A Probabilistic K-Rook Matching Perspective
topic Biomolecules
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2212.14041