MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Siyuan, Yu, Kai, Wang, Anna, Liu, Zicheng, Yu, Chang, Zhou, Jingbo, Yang, Qirong, Guo, Yucheng, Zhang, Xiaoming, Li, Stan Z.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917090165784576
author Li, Siyuan
Yu, Kai
Wang, Anna
Liu, Zicheng
Yu, Chang
Zhou, Jingbo
Yang, Qirong
Guo, Yucheng
Zhang, Xiaoming
Li, Stan Z.
author_facet Li, Siyuan
Yu, Kai
Wang, Anna
Liu, Zicheng
Yu, Chang
Zhou, Jingbo
Yang, Qirong
Guo, Yucheng
Zhang, Xiaoming
Li, Stan Z.
contents Modeling genomic sequences faces two unsolved challenges: the information density varies widely across different regions, while there is no clearly defined minimum vocabulary unit. Relying on either four primitive bases or independently designed DNA tokenizers, existing approaches with naive masked language modeling pre-training often fail to adapt to the varying complexities of genomic sequences. Leveraging Token Merging techniques, this paper introduces a hierarchical architecture that jointly optimizes a dynamic genomic tokenizer and latent Transformers with context-aware pre-training tasks. As for network structures, the tokenization module automatically chunks adjacent bases into words by stacking multiple layers of the differentiable token merging blocks with local-window constraints, then a Latent Encoder captures the global context of these merged words by full-attention blocks. Symmetrically employing a Latent Decoder and a Local Decoder, MergeDNA learns with two pre-training tasks: Merged Token Reconstruction simultaneously trains the dynamic tokenization module and adaptively filters important tokens, while Adaptive Masked Token Modeling learns to predict these filtered tokens to capture informative contents. Extensive experiments show that MergeDNA achieves superior performance on three popular DNA benchmarks and several multi-omics tasks with fine-tuning or zero-shot evaluation, outperforming typical tokenization methods and large-scale DNA foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14806
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
Li, Siyuan
Yu, Kai
Wang, Anna
Liu, Zicheng
Yu, Chang
Zhou, Jingbo
Yang, Qirong
Guo, Yucheng
Zhang, Xiaoming
Li, Stan Z.
Genomics
Artificial Intelligence
Machine Learning
Modeling genomic sequences faces two unsolved challenges: the information density varies widely across different regions, while there is no clearly defined minimum vocabulary unit. Relying on either four primitive bases or independently designed DNA tokenizers, existing approaches with naive masked language modeling pre-training often fail to adapt to the varying complexities of genomic sequences. Leveraging Token Merging techniques, this paper introduces a hierarchical architecture that jointly optimizes a dynamic genomic tokenizer and latent Transformers with context-aware pre-training tasks. As for network structures, the tokenization module automatically chunks adjacent bases into words by stacking multiple layers of the differentiable token merging blocks with local-window constraints, then a Latent Encoder captures the global context of these merged words by full-attention blocks. Symmetrically employing a Latent Decoder and a Local Decoder, MergeDNA learns with two pre-training tasks: Merged Token Reconstruction simultaneously trains the dynamic tokenization module and adaptively filters important tokens, while Adaptive Masked Token Modeling learns to predict these filtered tokens to capture informative contents. Extensive experiments show that MergeDNA achieves superior performance on three popular DNA benchmarks and several multi-omics tasks with fine-tuning or zero-shot evaluation, outperforming typical tokenization methods and large-scale DNA foundation models.
title MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
topic Genomics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.14806