Comprehensive benchmarking of large language models for RNA secondary structure prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zablocki, L. I., Bugnon, L. A., Gerard, M., Di Persia, L., Stegmayer, G., Milone, D. H.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913673625206784
author Zablocki, L. I.
Bugnon, L. A.
Gerard, M.
Di Persia, L.
Stegmayer, G.
Milone, D. H.
author_facet Zablocki, L. I.
Bugnon, L. A.
Gerard, M.
Di Persia, L.
Stegmayer, G.
Milone, D. H.
contents Inspired by the success of large language models (LLM) for DNA and proteins, several LLM for RNA have been developed recently. RNA-LLM uses large datasets of RNA sequences to learn, in a self-supervised way, how to represent each RNA base with a semantically rich numerical vector. This is done under the hypothesis that obtaining high-quality RNA representations can enhance data-costly downstream tasks. Among them, predicting the secondary structure is a fundamental task for uncovering RNA functional mechanisms. In this work we present a comprehensive experimental analysis of several pre-trained RNA-LLM, comparing them for the RNA secondary structure prediction task in an unified deep learning framework. The RNA-LLM were assessed with increasing generalization difficulty on benchmark datasets. Results showed that two LLM clearly outperform the other models, and revealed significant challenges for generalization in low-homology scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16212
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Comprehensive benchmarking of large language models for RNA secondary structure prediction
Zablocki, L. I.
Bugnon, L. A.
Gerard, M.
Di Persia, L.
Stegmayer, G.
Milone, D. H.
Artificial Intelligence
Machine Learning
Biomolecules
Inspired by the success of large language models (LLM) for DNA and proteins, several LLM for RNA have been developed recently. RNA-LLM uses large datasets of RNA sequences to learn, in a self-supervised way, how to represent each RNA base with a semantically rich numerical vector. This is done under the hypothesis that obtaining high-quality RNA representations can enhance data-costly downstream tasks. Among them, predicting the secondary structure is a fundamental task for uncovering RNA functional mechanisms. In this work we present a comprehensive experimental analysis of several pre-trained RNA-LLM, comparing them for the RNA secondary structure prediction task in an unified deep learning framework. The RNA-LLM were assessed with increasing generalization difficulty on benchmark datasets. Results showed that two LLM clearly outperform the other models, and revealed significant challenges for generalization in low-homology scenarios.
title Comprehensive benchmarking of large language models for RNA secondary structure prediction
topic Artificial Intelligence
Machine Learning
Biomolecules
url https://arxiv.org/abs/2410.16212