Count Bridges enable Modeling and Deconvolving Transcriptomic Data

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fishman, Nic, Gowri, Gokul, Kumar, Tanush, Lu, Jiaqi, de Bortoli, Valentin, Gootenberg, Jonathan S., Abudayyeh, Omar
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910041667272704
author Fishman, Nic
Gowri, Gokul
Kumar, Tanush
Lu, Jiaqi
de Bortoli, Valentin
Gootenberg, Jonathan S.
Abudayyeh, Omar
author_facet Fishman, Nic
Gowri, Gokul
Kumar, Tanush
Lu, Jiaqi
de Bortoli, Valentin
Gootenberg, Jonathan S.
Abudayyeh, Omar
contents Many modern biological assays, including RNA sequencing, yield integer-valued counts that reflect the number of molecules detected. These measurements are often not at the desired resolution: while the unit of interest is typically a single cell, many measurement technologies produce counts aggregated over sets of cells. Although recent generative frameworks such as diffusion and flow matching have been extended to non-Euclidean and discrete settings, it remains unclear how best to model integer-valued data or how to systematically deconvolve aggregated observations. We introduce Count Bridges, a stochastic bridge process on the integers that provides an exact, tractable analogue of diffusion-style models for count data, with closed-form conditionals for efficient training and sampling. We extend this framework to enable direct training from aggregated measurements via an Expectation-Maximization-style approach that treats unit-level counts as latent variables. We demonstrate state-of-the-art performance on integer distribution matching benchmarks, comparing against flow matching and discrete flow matching baselines across various metrics. We then apply Count Bridges to two large-scale problems in biology: modeling single-cell gene expression data at the nucleotide resolution, with applications to deconvolving bulk RNA-seq, and resolving multicellular spatial transcriptomic spots into single-cell count profiles. Our methods offer a principled foundation for generative modeling and deconvolution of biological count data across scales and modalities.
format Preprint
id arxiv_https___arxiv_org_abs_2603_04730
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Count Bridges enable Modeling and Deconvolving Transcriptomic Data
Fishman, Nic
Gowri, Gokul
Kumar, Tanush
Lu, Jiaqi
de Bortoli, Valentin
Gootenberg, Jonathan S.
Abudayyeh, Omar
Machine Learning
Many modern biological assays, including RNA sequencing, yield integer-valued counts that reflect the number of molecules detected. These measurements are often not at the desired resolution: while the unit of interest is typically a single cell, many measurement technologies produce counts aggregated over sets of cells. Although recent generative frameworks such as diffusion and flow matching have been extended to non-Euclidean and discrete settings, it remains unclear how best to model integer-valued data or how to systematically deconvolve aggregated observations. We introduce Count Bridges, a stochastic bridge process on the integers that provides an exact, tractable analogue of diffusion-style models for count data, with closed-form conditionals for efficient training and sampling. We extend this framework to enable direct training from aggregated measurements via an Expectation-Maximization-style approach that treats unit-level counts as latent variables. We demonstrate state-of-the-art performance on integer distribution matching benchmarks, comparing against flow matching and discrete flow matching baselines across various metrics. We then apply Count Bridges to two large-scale problems in biology: modeling single-cell gene expression data at the nucleotide resolution, with applications to deconvolving bulk RNA-seq, and resolving multicellular spatial transcriptomic spots into single-cell count profiles. Our methods offer a principled foundation for generative modeling and deconvolution of biological count data across scales and modalities.
title Count Bridges enable Modeling and Deconvolving Transcriptomic Data
topic Machine Learning
url https://arxiv.org/abs/2603.04730