PTCMIL: Multiple Instance Learning via Prompt Token Clustering for Whole Slide Image Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Beidi, Kim, SangMook, Chen, Hao, Zhou, Chen, Gao, Zu-hua, Wang, Gang, Li, Xiaoxiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908465498161152
author Zhao, Beidi
Kim, SangMook
Chen, Hao
Zhou, Chen
Gao, Zu-hua
Wang, Gang
Li, Xiaoxiao
author_facet Zhao, Beidi
Kim, SangMook
Chen, Hao
Zhou, Chen
Gao, Zu-hua
Wang, Gang
Li, Xiaoxiao
contents Multiple Instance Learning (MIL) has advanced WSI analysis but struggles with the complexity and heterogeneity of WSIs. Existing MIL methods face challenges in aggregating diverse patch information into robust WSI representations. While ViTs and clustering-based approaches show promise, they are computationally intensive and fail to capture task-specific and slide-specific variability. To address these limitations, we propose PTCMIL, a novel Prompt Token Clustering-based ViT for MIL aggregation. By introducing learnable prompt tokens into the ViT backbone, PTCMIL unifies clustering and prediction tasks in an end-to-end manner. It dynamically aligns clustering with downstream tasks, using projection-based clustering tailored to each WSI, reducing complexity while preserving patch heterogeneity. Through token merging and prototype-based pooling, PTCMIL efficiently captures task-relevant patterns. Extensive experiments on eight datasets demonstrate its superior performance in classification and survival analysis tasks, outperforming state-of-the-art methods. Systematic ablation studies confirm its robustness and strong interpretability. The code is released at https://github.com/ubc-tea/PTCMIL.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18848
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PTCMIL: Multiple Instance Learning via Prompt Token Clustering for Whole Slide Image Analysis
Zhao, Beidi
Kim, SangMook
Chen, Hao
Zhou, Chen
Gao, Zu-hua
Wang, Gang
Li, Xiaoxiao
Computer Vision and Pattern Recognition
Artificial Intelligence
Multiple Instance Learning (MIL) has advanced WSI analysis but struggles with the complexity and heterogeneity of WSIs. Existing MIL methods face challenges in aggregating diverse patch information into robust WSI representations. While ViTs and clustering-based approaches show promise, they are computationally intensive and fail to capture task-specific and slide-specific variability. To address these limitations, we propose PTCMIL, a novel Prompt Token Clustering-based ViT for MIL aggregation. By introducing learnable prompt tokens into the ViT backbone, PTCMIL unifies clustering and prediction tasks in an end-to-end manner. It dynamically aligns clustering with downstream tasks, using projection-based clustering tailored to each WSI, reducing complexity while preserving patch heterogeneity. Through token merging and prototype-based pooling, PTCMIL efficiently captures task-relevant patterns. Extensive experiments on eight datasets demonstrate its superior performance in classification and survival analysis tasks, outperforming state-of-the-art methods. Systematic ablation studies confirm its robustness and strong interpretability. The code is released at https://github.com/ubc-tea/PTCMIL.
title PTCMIL: Multiple Instance Learning via Prompt Token Clustering for Whole Slide Image Analysis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.18848