Single GPU Task Adaptation of Pathology Foundation Models for Whole Slide Image Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Neeraj, Nanda, Swaraj, Singi, Siddharth, Benhamida, Jamal, Kim, David, Chen, Jie-Fu, Momeni-Boroujeni, Amir, Goldgof, Gregory M., Campanella, Gabriele, Vanderbilt, Chad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916781192380416
author Kumar, Neeraj
Nanda, Swaraj
Singi, Siddharth
Benhamida, Jamal
Kim, David
Chen, Jie-Fu
Momeni-Boroujeni, Amir
Goldgof, Gregory M.
Campanella, Gabriele
Vanderbilt, Chad
author_facet Kumar, Neeraj
Nanda, Swaraj
Singi, Siddharth
Benhamida, Jamal
Kim, David
Chen, Jie-Fu
Momeni-Boroujeni, Amir
Goldgof, Gregory M.
Campanella, Gabriele
Vanderbilt, Chad
contents Pathology foundation models (PFMs) have emerged as powerful tools for analyzing whole slide images (WSIs). However, adapting these pretrained PFMs for specific clinical tasks presents considerable challenges, primarily due to the availability of only weak (WSI-level) labels for gigapixel images, necessitating multiple instance learning (MIL) paradigm for effective WSI analysis. This paper proposes a novel approach for single-GPU \textbf{T}ask \textbf{A}daptation of \textbf{PFM}s (TAPFM) that uses vision transformer (\vit) attention for MIL aggregation while optimizing both for feature representations and attention weights. The proposed approach maintains separate computational graphs for MIL aggregator and the PFM to create stable training dynamics that align with downstream task objectives during end-to-end adaptation. Evaluated on mutation prediction tasks for bladder cancer and lung adenocarcinoma across institutional and TCGA cohorts, TAPFM consistently outperforms conventional approaches, with H-Optimus-0 (TAPFM) outperforming the benchmarks. TAPFM effectively handles multi-label classification of actionable mutations as well. Thus, TAPFM makes adaptation of powerful pre-trained PFMs practical on standard hardware for various clinical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Single GPU Task Adaptation of Pathology Foundation Models for Whole Slide Image Analysis
Kumar, Neeraj
Nanda, Swaraj
Singi, Siddharth
Benhamida, Jamal
Kim, David
Chen, Jie-Fu
Momeni-Boroujeni, Amir
Goldgof, Gregory M.
Campanella, Gabriele
Vanderbilt, Chad
Computer Vision and Pattern Recognition
Pathology foundation models (PFMs) have emerged as powerful tools for analyzing whole slide images (WSIs). However, adapting these pretrained PFMs for specific clinical tasks presents considerable challenges, primarily due to the availability of only weak (WSI-level) labels for gigapixel images, necessitating multiple instance learning (MIL) paradigm for effective WSI analysis. This paper proposes a novel approach for single-GPU \textbf{T}ask \textbf{A}daptation of \textbf{PFM}s (TAPFM) that uses vision transformer (\vit) attention for MIL aggregation while optimizing both for feature representations and attention weights. The proposed approach maintains separate computational graphs for MIL aggregator and the PFM to create stable training dynamics that align with downstream task objectives during end-to-end adaptation. Evaluated on mutation prediction tasks for bladder cancer and lung adenocarcinoma across institutional and TCGA cohorts, TAPFM consistently outperforms conventional approaches, with H-Optimus-0 (TAPFM) outperforming the benchmarks. TAPFM effectively handles multi-label classification of actionable mutations as well. Thus, TAPFM makes adaptation of powerful pre-trained PFMs practical on standard hardware for various clinical applications.
title Single GPU Task Adaptation of Pathology Foundation Models for Whole Slide Image Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.05184