Latent Gene Diffusion for Spatial Transcriptomics Completion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cárdenas, Paula, Manrique, Leonardo, Vega, Daniela, Ruiz, Daniela, Arbeláez, Pablo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918134097641472
author Cárdenas, Paula
Manrique, Leonardo
Vega, Daniela
Ruiz, Daniela
Arbeláez, Pablo
author_facet Cárdenas, Paula
Manrique, Leonardo
Vega, Daniela
Ruiz, Daniela
Arbeláez, Pablo
contents Computer Vision has proven to be a powerful tool for analyzing Spatial Transcriptomics (ST) data. However, current models that predict spatially resolved gene expression from histopathology images suffer from significant limitations due to data dropout. Most existing approaches rely on single-cell RNA sequencing references, making them dependent on alignment quality and external datasets while also risking batch effects and inherited dropout. In this paper, we address these limitations by introducing LGDiST, the first reference-free latent gene diffusion model for ST data dropout. We show that LGDiST outperforms the previous state-of-the-art in gene expression completion, with an average Mean Squared Error that is 18% lower across 26 datasets. Furthermore, we demonstrate that completing ST data with LGDiST improves gene expression prediction performance on six state-of-the-art methods up to 10% in MSE. A key innovation of LGDiST is using context genes previously considered uninformative to build a rich and biologically meaningful genetic latent space. Our experiments show that removing key components of LGDiST, such as the context genes, the ST latent space, and the neighbor conditioning, leads to considerable drops in performance. These findings underscore that the full architecture of LGDiST achieves substantially better performance than any of its isolated components.
format Preprint
id arxiv_https___arxiv_org_abs_2509_01864
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Latent Gene Diffusion for Spatial Transcriptomics Completion
Cárdenas, Paula
Manrique, Leonardo
Vega, Daniela
Ruiz, Daniela
Arbeláez, Pablo
Computer Vision and Pattern Recognition
Computer Vision has proven to be a powerful tool for analyzing Spatial Transcriptomics (ST) data. However, current models that predict spatially resolved gene expression from histopathology images suffer from significant limitations due to data dropout. Most existing approaches rely on single-cell RNA sequencing references, making them dependent on alignment quality and external datasets while also risking batch effects and inherited dropout. In this paper, we address these limitations by introducing LGDiST, the first reference-free latent gene diffusion model for ST data dropout. We show that LGDiST outperforms the previous state-of-the-art in gene expression completion, with an average Mean Squared Error that is 18% lower across 26 datasets. Furthermore, we demonstrate that completing ST data with LGDiST improves gene expression prediction performance on six state-of-the-art methods up to 10% in MSE. A key innovation of LGDiST is using context genes previously considered uninformative to build a rich and biologically meaningful genetic latent space. Our experiments show that removing key components of LGDiST, such as the context genes, the ST latent space, and the neighbor conditioning, leads to considerable drops in performance. These findings underscore that the full architecture of LGDiST achieves substantially better performance than any of its isolated components.
title Latent Gene Diffusion for Spatial Transcriptomics Completion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.01864