PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Fengchun, Jiang, Songhan, Cai, Linghan, Wang, Ziyue, Zhang, Yongbing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912776997306368
author Liu, Fengchun
Jiang, Songhan
Cai, Linghan
Wang, Ziyue
Zhang, Yongbing
author_facet Liu, Fengchun
Jiang, Songhan
Cai, Linghan
Wang, Ziyue
Zhang, Yongbing
contents While Vision-Language Models (VLMs) have achieved notable progress in computational pathology (CPath), the gigapixel scale and spatial heterogeneity of Whole Slide Images (WSIs) continue to pose challenges for multimodal understanding. Existing alignment methods struggle to capture fine-grained correspondences between textual descriptions and visual cues across thousands of patches from a slide, compromising their performance on downstream tasks. In this paper, we propose PathFLIP (Pathology Fine-grained Language-Image Pretraining), a novel framework for holistic WSI interpretation. PathFLIP decomposes slide-level captions into region-level subcaptions and generates text-conditioned region embeddings to facilitate precise visual-language grounding. By harnessing Large Language Models (LLMs), PathFLIP can seamlessly follow diverse clinical instructions and adapt to varied diagnostic contexts. Furthermore, it exhibits versatile capabilities across multiple paradigms, efficiently handling slide-level classification and retrieval, fine-grained lesion localization, and instruction following. Extensive experiments demonstrate that PathFLIP outperforms existing large-scale pathological VLMs on four representative benchmarks while requiring significantly less training data, paving the way for fine-grained, instruction-aware WSI interpretation in clinical practice.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology
Liu, Fengchun
Jiang, Songhan
Cai, Linghan
Wang, Ziyue
Zhang, Yongbing
Computer Vision and Pattern Recognition
While Vision-Language Models (VLMs) have achieved notable progress in computational pathology (CPath), the gigapixel scale and spatial heterogeneity of Whole Slide Images (WSIs) continue to pose challenges for multimodal understanding. Existing alignment methods struggle to capture fine-grained correspondences between textual descriptions and visual cues across thousands of patches from a slide, compromising their performance on downstream tasks. In this paper, we propose PathFLIP (Pathology Fine-grained Language-Image Pretraining), a novel framework for holistic WSI interpretation. PathFLIP decomposes slide-level captions into region-level subcaptions and generates text-conditioned region embeddings to facilitate precise visual-language grounding. By harnessing Large Language Models (LLMs), PathFLIP can seamlessly follow diverse clinical instructions and adapt to varied diagnostic contexts. Furthermore, it exhibits versatile capabilities across multiple paradigms, efficiently handling slide-level classification and retrieval, fine-grained lesion localization, and instruction following. Extensive experiments demonstrate that PathFLIP outperforms existing large-scale pathological VLMs on four representative benchmarks while requiring significantly less training data, paving the way for fine-grained, instruction-aware WSI interpretation in clinical practice.
title PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.17621