Trace Reconstruction with Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Weindel, Franziska, Girsch, Michael, Heckel, Reinhard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915395309404160
author Weindel, Franziska
Girsch, Michael
Heckel, Reinhard
author_facet Weindel, Franziska
Girsch, Michael
Heckel, Reinhard
contents The general trace reconstruction problem seeks to recover an original sequence from its noisy copies independently corrupted by deletions, insertions, and substitutions. This problem arises in applications such as DNA data storage, a promising storage medium due to its high information density and longevity. However, errors introduced during DNA synthesis, storage, and sequencing require correction through algorithms and codes, with trace reconstruction often used as part of the data retrieval process. In this work, we propose TReconLM, which leverages language models trained on next-token prediction for trace reconstruction. We pretrain language models on synthetic data and fine-tune on real-world data to adapt to technology-specific error patterns. TReconLM outperforms state-of-the-art trace reconstruction algorithms, including prior deep learning approaches, recovering a substantially higher fraction of sequences without error.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12927
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Trace Reconstruction with Language Models
Weindel, Franziska
Girsch, Michael
Heckel, Reinhard
Machine Learning
Information Theory
The general trace reconstruction problem seeks to recover an original sequence from its noisy copies independently corrupted by deletions, insertions, and substitutions. This problem arises in applications such as DNA data storage, a promising storage medium due to its high information density and longevity. However, errors introduced during DNA synthesis, storage, and sequencing require correction through algorithms and codes, with trace reconstruction often used as part of the data retrieval process. In this work, we propose TReconLM, which leverages language models trained on next-token prediction for trace reconstruction. We pretrain language models on synthetic data and fine-tune on real-world data to adapt to technology-specific error patterns. TReconLM outperforms state-of-the-art trace reconstruction algorithms, including prior deep learning approaches, recovering a substantially higher fraction of sequences without error.
title Trace Reconstruction with Language Models
topic Machine Learning
Information Theory
url https://arxiv.org/abs/2507.12927