DNAMotifTokenizer: Towards Biologically Informed Tokenization of Genomic Sequences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Xiaoxiao, Wang, Zihan, Shang, Jingbo, Li, Yang E.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917162303619072
author Zhou, Xiaoxiao
Wang, Zihan
Shang, Jingbo
Li, Yang E.
author_facet Zhou, Xiaoxiao
Wang, Zihan
Shang, Jingbo
Li, Yang E.
contents DNA language models have advanced genomics, but their downstream performance varies widely due to differences in tokenization, pretraining data, and architecture. We argue that a major bottleneck lies in tokenizing sparse and unevenly distributed DNA sequence motifs, which are critical for accurate and interpretable models. To investigate, we systematically benchmark k-mer and Byte-Pair Encoding (BPE) tokenizers under controlled pretraining budget, evaluating across multiple downstream tasks from five datasets. We find that tokenizer choice induces task-specific trade-offs, and that vocabulary size and tokenizer training data strongly influence the biological knowledge captured. Notably, BPE tokenizers achieve strong performance when trained on smaller but biologically significant data. Building on these insights, we introduce DNAMotifTokenizer, which directly incorporates domain knowledge of DNA sequence motifs into the tokenization process. DNAMotifTokenizer consistently outperforms BPE across diverse benchmarks, demonstrating that knowledge-infused tokenization is crucial for learning powerful, interpretable, and generalizable genomic representations.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17126
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DNAMotifTokenizer: Towards Biologically Informed Tokenization of Genomic Sequences
Zhou, Xiaoxiao
Wang, Zihan
Shang, Jingbo
Li, Yang E.
Genomics
DNA language models have advanced genomics, but their downstream performance varies widely due to differences in tokenization, pretraining data, and architecture. We argue that a major bottleneck lies in tokenizing sparse and unevenly distributed DNA sequence motifs, which are critical for accurate and interpretable models. To investigate, we systematically benchmark k-mer and Byte-Pair Encoding (BPE) tokenizers under controlled pretraining budget, evaluating across multiple downstream tasks from five datasets. We find that tokenizer choice induces task-specific trade-offs, and that vocabulary size and tokenizer training data strongly influence the biological knowledge captured. Notably, BPE tokenizers achieve strong performance when trained on smaller but biologically significant data. Building on these insights, we introduce DNAMotifTokenizer, which directly incorporates domain knowledge of DNA sequence motifs into the tokenization process. DNAMotifTokenizer consistently outperforms BPE across diverse benchmarks, demonstrating that knowledge-infused tokenization is crucial for learning powerful, interpretable, and generalizable genomic representations.
title DNAMotifTokenizer: Towards Biologically Informed Tokenization of Genomic Sequences
topic Genomics
url https://arxiv.org/abs/2512.17126