Slide-Level Prompt Learning with Vision Language Models for Few-Shot Multiple Instance Learning in Histopathology

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tomar, Devavrat, Vray, Guillaume, Mahapatra, Dwarikanath, Roy, Sudipta, Thiran, Jean-Philippe, Bozorgtabar, Behzad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915208268611584
author Tomar, Devavrat
Vray, Guillaume
Mahapatra, Dwarikanath
Roy, Sudipta
Thiran, Jean-Philippe
Bozorgtabar, Behzad
author_facet Tomar, Devavrat
Vray, Guillaume
Mahapatra, Dwarikanath
Roy, Sudipta
Thiran, Jean-Philippe
Bozorgtabar, Behzad
contents In this paper, we address the challenge of few-shot classification in histopathology whole slide images (WSIs) by utilizing foundational vision-language models (VLMs) and slide-level prompt learning. Given the gigapixel scale of WSIs, conventional multiple instance learning (MIL) methods rely on aggregation functions to derive slide-level (bag-level) predictions from patch representations, which require extensive bag-level labels for training. In contrast, VLM-based approaches excel at aligning visual embeddings of patches with candidate class text prompts but lack essential pathological prior knowledge. Our method distinguishes itself by utilizing pathological prior knowledge from language models to identify crucial local tissue types (patches) for WSI classification, integrating this within a VLM-based MIL framework. Our approach effectively aligns patch images with tissue types, and we fine-tune our model via prompt learning using only a few labeled WSIs per category. Experimentation on real-world pathological WSI datasets and ablation studies highlight our method's superior performance over existing MIL- and VLM-based methods in few-shot WSI classification tasks. Our code is publicly available at https://github.com/LTS5/SLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17238
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Slide-Level Prompt Learning with Vision Language Models for Few-Shot Multiple Instance Learning in Histopathology
Tomar, Devavrat
Vray, Guillaume
Mahapatra, Dwarikanath
Roy, Sudipta
Thiran, Jean-Philippe
Bozorgtabar, Behzad
Computer Vision and Pattern Recognition
In this paper, we address the challenge of few-shot classification in histopathology whole slide images (WSIs) by utilizing foundational vision-language models (VLMs) and slide-level prompt learning. Given the gigapixel scale of WSIs, conventional multiple instance learning (MIL) methods rely on aggregation functions to derive slide-level (bag-level) predictions from patch representations, which require extensive bag-level labels for training. In contrast, VLM-based approaches excel at aligning visual embeddings of patches with candidate class text prompts but lack essential pathological prior knowledge. Our method distinguishes itself by utilizing pathological prior knowledge from language models to identify crucial local tissue types (patches) for WSI classification, integrating this within a VLM-based MIL framework. Our approach effectively aligns patch images with tissue types, and we fine-tune our model via prompt learning using only a few labeled WSIs per category. Experimentation on real-world pathological WSI datasets and ablation studies highlight our method's superior performance over existing MIL- and VLM-based methods in few-shot WSI classification tasks. Our code is publicly available at https://github.com/LTS5/SLIP.
title Slide-Level Prompt Learning with Vision Language Models for Few-Shot Multiple Instance Learning in Histopathology
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.17238