Finding low-complexity DNA sequences with longdust

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Heng, Li, Brian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914159663251456
author Li, Heng
Li, Brian
author_facet Li, Heng
Li, Brian
contents Motivation: Low-complexity (LC) DNA sequences are compositionally repetitive sequences that are often associated with spurious homologous matches and variant calling artifacts. While algorithms for identifying LC sequences exist, they either do not define complexity mathematically or are inefficient with long or variable context windows. Results: Longdust is a new algorithm that efficiently identifies long LC sequences including centromeric satellite and tandem repeats with moderately long motifs. It defines string complexity by statistically modeling the k-mer count distribution with the parameters: the k-mer length, the context window size and a threshold on complexity. Longdust exhibits high performance on real data and high consistency with existing methods. Availability and implementation: https://github.com/lh3/longdust
format Preprint
id arxiv_https___arxiv_org_abs_2509_07357
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Finding low-complexity DNA sequences with longdust
Li, Heng
Li, Brian
Genomics
Motivation: Low-complexity (LC) DNA sequences are compositionally repetitive sequences that are often associated with spurious homologous matches and variant calling artifacts. While algorithms for identifying LC sequences exist, they either do not define complexity mathematically or are inefficient with long or variable context windows. Results: Longdust is a new algorithm that efficiently identifies long LC sequences including centromeric satellite and tandem repeats with moderately long motifs. It defines string complexity by statistically modeling the k-mer count distribution with the parameters: the k-mer length, the context window size and a threshold on complexity. Longdust exhibits high performance on real data and high consistency with existing methods. Availability and implementation: https://github.com/lh3/longdust
title Finding low-complexity DNA sequences with longdust
topic Genomics
url https://arxiv.org/abs/2509.07357