annbatch unlocks terabyte-scale training of biological data in anndata

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gold, Ilan, Fischer, Felix, Arnoldt, Lucas, Wolf, F. Alexander, Theis, Fabian J.
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915911572652032
author Gold, Ilan
Fischer, Felix
Arnoldt, Lucas
Wolf, F. Alexander
Theis, Fabian J.
author_facet Gold, Ilan
Fischer, Felix
Arnoldt, Lucas
Wolf, F. Alexander
Theis, Fabian J.
contents The scale of biological datasets now routinely exceeds system memory, making data access rather than model computation the primary bottleneck in training machine-learning models. This bottleneck is particularly acute in biology, where widely used community data formats must support heterogeneous metadata, sparse and dense assays, and downstream analysis within established computational ecosystems. Here we present annbatch, a mini-batch loader native to anndata that enables out-of-core training directly on disk-backed datasets. Across single-cell transcriptomics, microscopy and whole-genome sequencing benchmarks, annbatch increases loading throughput by up to an order of magnitude and shortens training from days to hours, while remaining fully compatible with the scverse ecosystem. Annbatch establishes a practical data-loading infrastructure for scalable biological AI, allowing increasingly large and diverse datasets to be used without abandoning standard biological data formats. Github: https://github.com/scverse/annbatch
format Preprint
id arxiv_https___arxiv_org_abs_2604_01949
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle annbatch unlocks terabyte-scale training of biological data in anndata
Gold, Ilan
Fischer, Felix
Arnoldt, Lucas
Wolf, F. Alexander
Theis, Fabian J.
Machine Learning
Genomics
The scale of biological datasets now routinely exceeds system memory, making data access rather than model computation the primary bottleneck in training machine-learning models. This bottleneck is particularly acute in biology, where widely used community data formats must support heterogeneous metadata, sparse and dense assays, and downstream analysis within established computational ecosystems. Here we present annbatch, a mini-batch loader native to anndata that enables out-of-core training directly on disk-backed datasets. Across single-cell transcriptomics, microscopy and whole-genome sequencing benchmarks, annbatch increases loading throughput by up to an order of magnitude and shortens training from days to hours, while remaining fully compatible with the scverse ecosystem. Annbatch establishes a practical data-loading infrastructure for scalable biological AI, allowing increasingly large and diverse datasets to be used without abandoning standard biological data formats. Github: https://github.com/scverse/annbatch
title annbatch unlocks terabyte-scale training of biological data in anndata
topic Machine Learning
Genomics
url https://arxiv.org/abs/2604.01949