NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhongmin, Ma, Runze, Tan, Jiahao, Tan, Chengzi, Zheng, Shuangjia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908629352841216
author Li, Zhongmin
Ma, Runze
Tan, Jiahao
Tan, Chengzi
Zheng, Shuangjia
author_facet Li, Zhongmin
Ma, Runze
Tan, Jiahao
Tan, Chengzi
Zheng, Shuangjia
contents Nucleotide sequence variation can induce significant shifts in functional fitness. Recent nucleotide foundation models promise to predict such fitness effects directly from sequence, yet heterogeneous datasets and inconsistent preprocessing make it difficult to compare methods fairly across DNA and RNA families. Here we introduce NABench, a large-scale, systematic benchmark for nucleic acid fitness prediction. NABench aggregates 162 high-throughput assays and curates 2.6 million mutated sequences spanning diverse DNA and RNA families, with standardized splits and rich metadata. We show that NABench surpasses prior nucleotide fitness benchmarks in scale, diversity, and data quality. Under a unified evaluation suite, we rigorously assess 29 representative foundation models across zero-shot, few-shot prediction, transfer learning, and supervised settings. The results quantify performance heterogeneity across tasks and nucleic-acid types, demonstrating clear strengths and failure modes for different modeling choices and establishing strong, reproducible baselines. We release NABench to advance nucleic acid modeling, supporting downstream applications in RNA/DNA design, synthetic biology, and biochemistry. Our code is available at https://github.com/mrzzmrzz/NABench.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02888
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction
Li, Zhongmin
Ma, Runze
Tan, Jiahao
Tan, Chengzi
Zheng, Shuangjia
Genomics
Artificial Intelligence
Nucleotide sequence variation can induce significant shifts in functional fitness. Recent nucleotide foundation models promise to predict such fitness effects directly from sequence, yet heterogeneous datasets and inconsistent preprocessing make it difficult to compare methods fairly across DNA and RNA families. Here we introduce NABench, a large-scale, systematic benchmark for nucleic acid fitness prediction. NABench aggregates 162 high-throughput assays and curates 2.6 million mutated sequences spanning diverse DNA and RNA families, with standardized splits and rich metadata. We show that NABench surpasses prior nucleotide fitness benchmarks in scale, diversity, and data quality. Under a unified evaluation suite, we rigorously assess 29 representative foundation models across zero-shot, few-shot prediction, transfer learning, and supervised settings. The results quantify performance heterogeneity across tasks and nucleic-acid types, demonstrating clear strengths and failure modes for different modeling choices and establishing strong, reproducible baselines. We release NABench to advance nucleic acid modeling, supporting downstream applications in RNA/DNA design, synthetic biology, and biochemistry. Our code is available at https://github.com/mrzzmrzz/NABench.
title NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction
topic Genomics
Artificial Intelligence
url https://arxiv.org/abs/2511.02888