Permutation Recovery Problem against Deletion Errors for DNA Data Storage

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singhvi, Shubhransh, Gupta, Charchit, Boruchovsky, Avital, Goldberg, Yuval, Kiah, Han Mao, Yaakobi, Eitan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929286643974144
author Singhvi, Shubhransh
Gupta, Charchit
Boruchovsky, Avital
Goldberg, Yuval
Kiah, Han Mao
Yaakobi, Eitan
author_facet Singhvi, Shubhransh
Gupta, Charchit
Boruchovsky, Avital
Goldberg, Yuval
Kiah, Han Mao
Yaakobi, Eitan
contents Owing to its immense storage density and durability, DNA has emerged as a promising storage medium. However, due to technological constraints, data can only be written onto many short DNA molecules called data blocks that are stored in an unordered way. To handle the unordered nature of DNA data storage systems, a unique address is typically prepended to each data block to form a DNA strand. However, DNA storage systems are prone to errors and generate multiple noisy copies of each strand called DNA reads. Thus, we study the permutation recovery problem against deletions errors for DNA data storage. The permutation recovery problem for DNA data storage requires one to reconstruct the addresses or in other words to uniquely identify the noisy reads. By successfully reconstructing the addresses, one can essentially determine the correct order of the data blocks, effectively solving the clustering problem. We first show that we can almost surely identify all the noisy reads under certain mild assumptions. We then propose a permutation recovery procedure and analyze its complexity.
format Preprint
id arxiv_https___arxiv_org_abs_2403_15827
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Permutation Recovery Problem against Deletion Errors for DNA Data Storage
Singhvi, Shubhransh
Gupta, Charchit
Boruchovsky, Avital
Goldberg, Yuval
Kiah, Han Mao
Yaakobi, Eitan
Information Theory
Owing to its immense storage density and durability, DNA has emerged as a promising storage medium. However, due to technological constraints, data can only be written onto many short DNA molecules called data blocks that are stored in an unordered way. To handle the unordered nature of DNA data storage systems, a unique address is typically prepended to each data block to form a DNA strand. However, DNA storage systems are prone to errors and generate multiple noisy copies of each strand called DNA reads. Thus, we study the permutation recovery problem against deletions errors for DNA data storage. The permutation recovery problem for DNA data storage requires one to reconstruct the addresses or in other words to uniquely identify the noisy reads. By successfully reconstructing the addresses, one can essentially determine the correct order of the data blocks, effectively solving the clustering problem. We first show that we can almost surely identify all the noisy reads under certain mild assumptions. We then propose a permutation recovery procedure and analyze its complexity.
title Permutation Recovery Problem against Deletion Errors for DNA Data Storage
topic Information Theory
url https://arxiv.org/abs/2403.15827