MGPATH: Vision-Language Model with Multi-Granular Prompt Learning for Few-Shot WSI Classification

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nguyen, Anh-Tien, Nguyen, Duy Minh Ho, Diep, Nghiem Tuong, Nguyen, Trung Quoc, Ho, Nhat, Metsch, Jacqueline Michelle, Maurer, Miriam Cindy, Sonntag, Daniel, Bohnenberger, Hanibal, Hauschild, Anne-Christin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909881226756096
author Nguyen, Anh-Tien
Nguyen, Duy Minh Ho
Diep, Nghiem Tuong
Nguyen, Trung Quoc
Ho, Nhat
Metsch, Jacqueline Michelle
Maurer, Miriam Cindy
Sonntag, Daniel
Bohnenberger, Hanibal
Hauschild, Anne-Christin
author_facet Nguyen, Anh-Tien
Nguyen, Duy Minh Ho
Diep, Nghiem Tuong
Nguyen, Trung Quoc
Ho, Nhat
Metsch, Jacqueline Michelle
Maurer, Miriam Cindy
Sonntag, Daniel
Bohnenberger, Hanibal
Hauschild, Anne-Christin
contents Whole slide pathology image classification presents challenges due to gigapixel image sizes and limited annotation labels, hindering model generalization. This paper introduces a prompt learning method to adapt large vision-language models for few-shot pathology classification. We first extend the Prov-GigaPath vision foundation model, pre-trained on 1.3 billion pathology image tiles, into a vision-language model by adding adaptors and aligning it with medical text encoders via contrastive learning on 923K image-text pairs. The model is then used to extract visual features and text embeddings from few-shot annotations and fine-tunes with learnable prompt embeddings. Unlike prior methods that combine prompts with frozen features using prefix embeddings or self-attention, we propose multi-granular attention that compares interactions between learnable prompts with individual image patches and groups of them. This approach improves the model's ability to capture both fine-grained details and broader context, enhancing its recognition of complex patterns across sub-regions. To further improve accuracy, we leverage (unbalanced) optimal transport-based visual-text distance to secure model robustness by mitigating perturbations that might occur during the data augmentation process. Empirical experiments on lung, kidney, and breast pathology modalities validate the effectiveness of our approach; thereby, we surpass several of the latest competitors and consistently improve performance across diverse architectures, including CLIP, PLIP, and Prov-GigaPath integrated PLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07409
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MGPATH: Vision-Language Model with Multi-Granular Prompt Learning for Few-Shot WSI Classification
Nguyen, Anh-Tien
Nguyen, Duy Minh Ho
Diep, Nghiem Tuong
Nguyen, Trung Quoc
Ho, Nhat
Metsch, Jacqueline Michelle
Maurer, Miriam Cindy
Sonntag, Daniel
Bohnenberger, Hanibal
Hauschild, Anne-Christin
Computer Vision and Pattern Recognition
Machine Learning
Whole slide pathology image classification presents challenges due to gigapixel image sizes and limited annotation labels, hindering model generalization. This paper introduces a prompt learning method to adapt large vision-language models for few-shot pathology classification. We first extend the Prov-GigaPath vision foundation model, pre-trained on 1.3 billion pathology image tiles, into a vision-language model by adding adaptors and aligning it with medical text encoders via contrastive learning on 923K image-text pairs. The model is then used to extract visual features and text embeddings from few-shot annotations and fine-tunes with learnable prompt embeddings. Unlike prior methods that combine prompts with frozen features using prefix embeddings or self-attention, we propose multi-granular attention that compares interactions between learnable prompts with individual image patches and groups of them. This approach improves the model's ability to capture both fine-grained details and broader context, enhancing its recognition of complex patterns across sub-regions. To further improve accuracy, we leverage (unbalanced) optimal transport-based visual-text distance to secure model robustness by mitigating perturbations that might occur during the data augmentation process. Empirical experiments on lung, kidney, and breast pathology modalities validate the effectiveness of our approach; thereby, we surpass several of the latest competitors and consistently improve performance across diverse architectures, including CLIP, PLIP, and Prov-GigaPath integrated PLIP.
title MGPATH: Vision-Language Model with Multi-Granular Prompt Learning for Few-Shot WSI Classification
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2502.07409