TrinityDNA: A Bio-Inspired Foundational Model for Efficient Long-Sequence DNA Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Qirong, Guo, Yucheng, Liu, Zicheng, Yang, Yujie, Yin, Qijin, Li, Siyuan, Ji, Shaomin, Chao, Linlin, Zhang, Xiaoming, Li, Stan Z.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911491237609472
author Yang, Qirong
Guo, Yucheng
Liu, Zicheng
Yang, Yujie
Yin, Qijin
Li, Siyuan
Ji, Shaomin
Chao, Linlin
Zhang, Xiaoming
Li, Stan Z.
author_facet Yang, Qirong
Guo, Yucheng
Liu, Zicheng
Yang, Yujie
Yin, Qijin
Li, Siyuan
Ji, Shaomin
Chao, Linlin
Zhang, Xiaoming
Li, Stan Z.
contents The modeling of genomic sequences presents unique challenges due to their length and structural complexity. Traditional sequence models struggle to capture long-range dependencies and biological features inherent in DNA. In this work, we propose TrinityDNA, a novel DNA foundational model designed to address these challenges. The model integrates biologically informed components, including Groove Fusion for capturing DNA's structural features and Gated Reverse Complement (GRC) to handle the inherent symmetry of DNA sequences. Additionally, we introduce a multi-scale attention mechanism that allows the model to attend to varying levels of sequence dependencies, and an evolutionary training strategy that progressively adapts the model to both prokaryotic and eukaryotic genomes. TrinityDNA provides a more accurate and efficient approach to genomic sequence modeling, offering significant improvements in gene function prediction, regulatory mechanism discovery, and other genomics applications. Our model bridges the gap between machine learning techniques and biological insights, paving the way for more effective analysis of genomic data. Additionally, we introduced a new DNA long-sequence CDS annotation benchmark to make evaluations more comprehensive and oriented toward practical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19229
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TrinityDNA: A Bio-Inspired Foundational Model for Efficient Long-Sequence DNA Modeling
Yang, Qirong
Guo, Yucheng
Liu, Zicheng
Yang, Yujie
Yin, Qijin
Li, Siyuan
Ji, Shaomin
Chao, Linlin
Zhang, Xiaoming
Li, Stan Z.
Computational Engineering, Finance, and Science
Genomics
The modeling of genomic sequences presents unique challenges due to their length and structural complexity. Traditional sequence models struggle to capture long-range dependencies and biological features inherent in DNA. In this work, we propose TrinityDNA, a novel DNA foundational model designed to address these challenges. The model integrates biologically informed components, including Groove Fusion for capturing DNA's structural features and Gated Reverse Complement (GRC) to handle the inherent symmetry of DNA sequences. Additionally, we introduce a multi-scale attention mechanism that allows the model to attend to varying levels of sequence dependencies, and an evolutionary training strategy that progressively adapts the model to both prokaryotic and eukaryotic genomes. TrinityDNA provides a more accurate and efficient approach to genomic sequence modeling, offering significant improvements in gene function prediction, regulatory mechanism discovery, and other genomics applications. Our model bridges the gap between machine learning techniques and biological insights, paving the way for more effective analysis of genomic data. Additionally, we introduced a new DNA long-sequence CDS annotation benchmark to make evaluations more comprehensive and oriented toward practical applications.
title TrinityDNA: A Bio-Inspired Foundational Model for Efficient Long-Sequence DNA Modeling
topic Computational Engineering, Finance, and Science
Genomics
url https://arxiv.org/abs/2507.19229