MultiST: A Cross-Attention-Based Multimodal Model for Spatial Transcriptomic

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Wei, Ly, Quoc-Toan, Yu, Chong, Bai, Jun
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911386678853632
author Wang, Wei
Ly, Quoc-Toan
Yu, Chong
Bai, Jun
author_facet Wang, Wei
Ly, Quoc-Toan
Yu, Chong
Bai, Jun
contents Spatial transcriptomics (ST) enables transcriptome-wide profiling while preserving the spatial context of tissues, offering unprecedented opportunities to study tissue organization and cell-cell interactions in situ. Despite recent advances, existing methods often lack effective integration of histological morphology with molecular profiles, relying on shallow fusion strategies or omitting tissue images altogether, which limits their ability to resolve ambiguous spatial domain boundaries. To address this challenge, we propose MultiST, a unified multimodal framework that jointly models spatial topology, gene expression, and tissue morphology through cross-attention-based fusion. MultiST employs graph-based gene encoders with adversarial alignment to learn robust spatial representations, while integrating color-normalized histological features to capture molecular-morphological dependencies and refine domain boundaries. We evaluated the proposed method on 13 diverse ST datasets spanning two organs, including human brain cortex and breast cancer tissue. MultiST yields spatial domains with clearer and more coherent boundaries than existing methods, leading to more stable pseudotime trajectories and more biologically interpretable cell-cell interaction patterns. The MultiST framework and source code are available at https://github.com/LabJunBMI/MultiST.git.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13331
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MultiST: A Cross-Attention-Based Multimodal Model for Spatial Transcriptomic
Wang, Wei
Ly, Quoc-Toan
Yu, Chong
Bai, Jun
Computer Vision and Pattern Recognition
Machine Learning
Spatial transcriptomics (ST) enables transcriptome-wide profiling while preserving the spatial context of tissues, offering unprecedented opportunities to study tissue organization and cell-cell interactions in situ. Despite recent advances, existing methods often lack effective integration of histological morphology with molecular profiles, relying on shallow fusion strategies or omitting tissue images altogether, which limits their ability to resolve ambiguous spatial domain boundaries. To address this challenge, we propose MultiST, a unified multimodal framework that jointly models spatial topology, gene expression, and tissue morphology through cross-attention-based fusion. MultiST employs graph-based gene encoders with adversarial alignment to learn robust spatial representations, while integrating color-normalized histological features to capture molecular-morphological dependencies and refine domain boundaries. We evaluated the proposed method on 13 diverse ST datasets spanning two organs, including human brain cortex and breast cancer tissue. MultiST yields spatial domains with clearer and more coherent boundaries than existing methods, leading to more stable pseudotime trajectories and more biologically interpretable cell-cell interaction patterns. The MultiST framework and source code are available at https://github.com/LabJunBMI/MultiST.git.
title MultiST: A Cross-Attention-Based Multimodal Model for Spatial Transcriptomic
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2601.13331