PathoHR: Hierarchical Reasoning for Vision-Language Models in Pathology

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yating, Huang, Ziyan, Xiang, Lintao, Yang, Qijun, Yin, Hujun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916978691670016
author Huang, Yating
Huang, Ziyan
Xiang, Lintao
Yang, Qijun
Yin, Hujun
author_facet Huang, Yating
Huang, Ziyan
Xiang, Lintao
Yang, Qijun
Yin, Hujun
contents Accurate analysis of pathological images is essential for automated tumor diagnosis but remains challenging due to high structural similarity and subtle morphological variations in tissue images. Current vision-language (VL) models often struggle to capture the complex reasoning required for interpreting structured pathological reports. To address these limitations, we propose PathoHR-Bench, a novel benchmark designed to evaluate VL models' abilities in hierarchical semantic understanding and compositional reasoning within the pathology domain. Results of this benchmark reveal that existing VL models fail to effectively model intricate cross-modal relationships, hence limiting their applicability in clinical setting. To overcome this, we further introduce a pathology-specific VL training scheme that generates enhanced and perturbed samples for multimodal contrastive learning. Experimental evaluations demonstrate that our approach achieves state-of-the-art performance on PathoHR-Bench and six additional pathology datasets, highlighting its effectiveness in fine-grained pathology representation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06105
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PathoHR: Hierarchical Reasoning for Vision-Language Models in Pathology
Huang, Yating
Huang, Ziyan
Xiang, Lintao
Yang, Qijun
Yin, Hujun
Computer Vision and Pattern Recognition
Accurate analysis of pathological images is essential for automated tumor diagnosis but remains challenging due to high structural similarity and subtle morphological variations in tissue images. Current vision-language (VL) models often struggle to capture the complex reasoning required for interpreting structured pathological reports. To address these limitations, we propose PathoHR-Bench, a novel benchmark designed to evaluate VL models' abilities in hierarchical semantic understanding and compositional reasoning within the pathology domain. Results of this benchmark reveal that existing VL models fail to effectively model intricate cross-modal relationships, hence limiting their applicability in clinical setting. To overcome this, we further introduce a pathology-specific VL training scheme that generates enhanced and perturbed samples for multimodal contrastive learning. Experimental evaluations demonstrate that our approach achieves state-of-the-art performance on PathoHR-Bench and six additional pathology datasets, highlighting its effectiveness in fine-grained pathology representation.
title PathoHR: Hierarchical Reasoning for Vision-Language Models in Pathology
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.06105