RiNALMo: General-Purpose RNA Language Models Can Generalize Well on Structure Prediction Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Penić, Rafael Josip, Vlašić, Tin, Huber, Roland G., Wan, Yue, Šikić, Mile
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912477561749504
author Penić, Rafael Josip
Vlašić, Tin
Huber, Roland G.
Wan, Yue
Šikić, Mile
author_facet Penić, Rafael Josip
Vlašić, Tin
Huber, Roland G.
Wan, Yue
Šikić, Mile
contents While RNA has recently been recognized as an interesting small-molecule drug target, many challenges remain to be addressed before we take full advantage of it. This emphasizes the necessity to improve our understanding of its structures and functions. Over the years, sequencing technologies have produced an enormous amount of unlabeled RNA data, which hides a huge potential. Motivated by the successes of protein language models, we introduce RiboNucleic Acid Language Model (RiNALMo) to unveil the hidden code of RNA. RiNALMo is the largest RNA language model to date, with 650M parameters pre-trained on 36M non-coding RNA sequences from several databases. It can extract hidden knowledge and capture the underlying structure information implicitly embedded within the RNA sequences. RiNALMo achieves state-of-the-art results on several downstream tasks. Notably, we show that its generalization capabilities overcome the inability of other deep learning methods for secondary structure prediction to generalize on unseen RNA families.
format Preprint
id arxiv_https___arxiv_org_abs_2403_00043
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RiNALMo: General-Purpose RNA Language Models Can Generalize Well on Structure Prediction Tasks
Penić, Rafael Josip
Vlašić, Tin
Huber, Roland G.
Wan, Yue
Šikić, Mile
Biomolecules
Machine Learning
While RNA has recently been recognized as an interesting small-molecule drug target, many challenges remain to be addressed before we take full advantage of it. This emphasizes the necessity to improve our understanding of its structures and functions. Over the years, sequencing technologies have produced an enormous amount of unlabeled RNA data, which hides a huge potential. Motivated by the successes of protein language models, we introduce RiboNucleic Acid Language Model (RiNALMo) to unveil the hidden code of RNA. RiNALMo is the largest RNA language model to date, with 650M parameters pre-trained on 36M non-coding RNA sequences from several databases. It can extract hidden knowledge and capture the underlying structure information implicitly embedded within the RNA sequences. RiNALMo achieves state-of-the-art results on several downstream tasks. Notably, we show that its generalization capabilities overcome the inability of other deep learning methods for secondary structure prediction to generalize on unseen RNA families.
title RiNALMo: General-Purpose RNA Language Models Can Generalize Well on Structure Prediction Tasks
topic Biomolecules
Machine Learning
url https://arxiv.org/abs/2403.00043