Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hongyi, Zhu, Zhengjie, Ma, Jiabo, Wang, Fang, Shi, Yue, Luo, Bo, Wang, Jili, Cai, Qiuyu, Zhang, Xiuming, Chen, Yen-Wei, Lin, Lanfen, Chen, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911234709782528
author Wang, Hongyi
Zhu, Zhengjie
Ma, Jiabo
Wang, Fang
Shi, Yue
Luo, Bo
Wang, Jili
Cai, Qiuyu
Zhang, Xiuming
Chen, Yen-Wei
Lin, Lanfen
Chen, Hao
author_facet Wang, Hongyi
Zhu, Zhengjie
Ma, Jiabo
Wang, Fang
Shi, Yue
Luo, Bo
Wang, Jili
Cai, Qiuyu
Zhang, Xiuming
Chen, Yen-Wei
Lin, Lanfen
Chen, Hao
contents The rapid digitization of histopathology slides has opened up new possibilities for computational tools in clinical and research workflows. Among these, content-based slide retrieval stands out, enabling pathologists to identify morphologically and semantically similar cases, thereby supporting precise diagnoses, enhancing consistency across observers, and assisting example-based education. However, effective retrieval of whole slide images (WSIs) remains challenging due to their gigapixel scale and the difficulty of capturing subtle semantic differences amid abundant irrelevant content. To overcome these challenges, we present PathSearch, a retrieval framework that unifies fine-grained attentive mosaic representations with global-wise slide embeddings aligned through vision-language contrastive learning. Trained on a corpus of 6,926 slide-report pairs, PathSearch captures both fine-grained morphological cues and high-level semantic patterns to enable accurate and flexible retrieval. The framework supports two key functionalities: (1) mosaic-based image-to-image retrieval, ensuring accurate and efficient slide research; and (2) multi-modal retrieval, where text queries can directly retrieve relevant slides. PathSearch was rigorously evaluated on four public pathology datasets and three in-house cohorts, covering tasks including anatomical site retrieval, tumor subtyping, tumor vs. non-tumor discrimination, and grading across diverse organs such as breast, lung, kidney, liver, and stomach. External results show that PathSearch outperforms traditional image-to-image retrieval frameworks. A multi-center reader study further demonstrates that PathSearch improves diagnostic accuracy, boosts confidence, and enhances inter-observer agreement among pathologists in real clinical scenarios. These results establish PathSearch as a scalable and generalizable retrieval solution for digital pathology.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23224
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
Wang, Hongyi
Zhu, Zhengjie
Ma, Jiabo
Wang, Fang
Shi, Yue
Luo, Bo
Wang, Jili
Cai, Qiuyu
Zhang, Xiuming
Chen, Yen-Wei
Lin, Lanfen
Chen, Hao
Computer Vision and Pattern Recognition
Information Retrieval
The rapid digitization of histopathology slides has opened up new possibilities for computational tools in clinical and research workflows. Among these, content-based slide retrieval stands out, enabling pathologists to identify morphologically and semantically similar cases, thereby supporting precise diagnoses, enhancing consistency across observers, and assisting example-based education. However, effective retrieval of whole slide images (WSIs) remains challenging due to their gigapixel scale and the difficulty of capturing subtle semantic differences amid abundant irrelevant content. To overcome these challenges, we present PathSearch, a retrieval framework that unifies fine-grained attentive mosaic representations with global-wise slide embeddings aligned through vision-language contrastive learning. Trained on a corpus of 6,926 slide-report pairs, PathSearch captures both fine-grained morphological cues and high-level semantic patterns to enable accurate and flexible retrieval. The framework supports two key functionalities: (1) mosaic-based image-to-image retrieval, ensuring accurate and efficient slide research; and (2) multi-modal retrieval, where text queries can directly retrieve relevant slides. PathSearch was rigorously evaluated on four public pathology datasets and three in-house cohorts, covering tasks including anatomical site retrieval, tumor subtyping, tumor vs. non-tumor discrimination, and grading across diverse organs such as breast, lung, kidney, liver, and stomach. External results show that PathSearch outperforms traditional image-to-image retrieval frameworks. A multi-center reader study further demonstrates that PathSearch improves diagnostic accuracy, boosts confidence, and enhances inter-observer agreement among pathologists in real clinical scenarios. These results establish PathSearch as a scalable and generalizable retrieval solution for digital pathology.
title Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
topic Computer Vision and Pattern Recognition
Information Retrieval
url https://arxiv.org/abs/2510.23224