Saved in:
Bibliographic Details
Main Authors: Tkachenko, Adrian, Salem, Sepehr, Adeniyi, Ayotomiwa Ezekiel, Bingol, Zulal, Uddin, Mohammed Nayeem, Prasanna, Akshat, Zelikovsky, Alexander, Mangul, Serghei, Alkan, Can, Alser, Mohammed
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2601.17184
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911395822436352
author Tkachenko, Adrian
Salem, Sepehr
Adeniyi, Ayotomiwa Ezekiel
Bingol, Zulal
Uddin, Mohammed Nayeem
Prasanna, Akshat
Zelikovsky, Alexander
Mangul, Serghei
Alkan, Can
Alser, Mohammed
author_facet Tkachenko, Adrian
Salem, Sepehr
Adeniyi, Ayotomiwa Ezekiel
Bingol, Zulal
Uddin, Mohammed Nayeem
Prasanna, Akshat
Zelikovsky, Alexander
Mangul, Serghei
Alkan, Can
Alser, Mohammed
contents Motivation: High-throughput sequencing (HTS) enables population-scale genomics but generates massive datasets, creating bottlenecks in storage, transfer, and analysis. FASTQ, the standard format for over two decades, stores one byte per base and one byte per quality score, leading to inefficient I/O, high storage costs, and redundancy. Existing compression tools can mitigate some issues, but often introduce costly decompression or complex dependency issues. Results: We introduce FASTR, a lossless, computation-native successor to FASTQ that encodes each nucleotide together with its base quality score into a single 8-bit value. FASTR reduces file size by at least 2x while remaining fully reversible and directly usable for downstream analyses. Applying general-purpose compression tools on FASTR consistently yields higher compression ratios, 2.47, 3.64, and 4.8x faster compression, and 2.34, 1.96, 1.75x faster decompression than on FASTQ across Illumina, HiFi, and ONT reads. FASTR is machine-learning-ready, allowing reads to be consumed directly as numerical vectors or image-like representations. We provide a highly parallel software ecosystem for FASTQ-FASTR conversion and show that FASTR integrates with existing tools, such as minimap2, with minimal interface changes and no performance overhead. By eliminating decompression costs and reducing data movement, FASTR lays the foundation for scalable genomics analyses and real-time sequencing workflows. Availability and Implementation: https://github.com/ALSER-Lab/FASTR
format Preprint
id arxiv_https___arxiv_org_abs_2601_17184
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FASTR: Reimagining FASTQ via Compact Image-inspired Representation
Tkachenko, Adrian
Salem, Sepehr
Adeniyi, Ayotomiwa Ezekiel
Bingol, Zulal
Uddin, Mohammed Nayeem
Prasanna, Akshat
Zelikovsky, Alexander
Mangul, Serghei
Alkan, Can
Alser, Mohammed
Genomics
Machine Learning
Motivation: High-throughput sequencing (HTS) enables population-scale genomics but generates massive datasets, creating bottlenecks in storage, transfer, and analysis. FASTQ, the standard format for over two decades, stores one byte per base and one byte per quality score, leading to inefficient I/O, high storage costs, and redundancy. Existing compression tools can mitigate some issues, but often introduce costly decompression or complex dependency issues. Results: We introduce FASTR, a lossless, computation-native successor to FASTQ that encodes each nucleotide together with its base quality score into a single 8-bit value. FASTR reduces file size by at least 2x while remaining fully reversible and directly usable for downstream analyses. Applying general-purpose compression tools on FASTR consistently yields higher compression ratios, 2.47, 3.64, and 4.8x faster compression, and 2.34, 1.96, 1.75x faster decompression than on FASTQ across Illumina, HiFi, and ONT reads. FASTR is machine-learning-ready, allowing reads to be consumed directly as numerical vectors or image-like representations. We provide a highly parallel software ecosystem for FASTQ-FASTR conversion and show that FASTR integrates with existing tools, such as minimap2, with minimal interface changes and no performance overhead. By eliminating decompression costs and reducing data movement, FASTR lays the foundation for scalable genomics analyses and real-time sequencing workflows. Availability and Implementation: https://github.com/ALSER-Lab/FASTR
title FASTR: Reimagining FASTQ via Compact Image-inspired Representation
topic Genomics
Machine Learning
url https://arxiv.org/abs/2601.17184