PathoTune: Adapting Visual Foundation Model to Pathological Specialists

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Jiaxuan, Yan, Fang, Zhang, Xiaofan, Gao, Yue, Zhang, Shaoting
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910526193270784
author Lu, Jiaxuan
Yan, Fang
Zhang, Xiaofan
Gao, Yue
Zhang, Shaoting
author_facet Lu, Jiaxuan
Yan, Fang
Zhang, Xiaofan
Gao, Yue
Zhang, Shaoting
contents As natural image understanding moves towards the pretrain-finetune era, research in pathology imaging is concurrently evolving. Despite the predominant focus on pretraining pathological foundation models, how to adapt foundation models to downstream tasks is little explored. For downstream adaptation, we propose the existence of two domain gaps, i.e., the Foundation-Task Gap and the Task-Instance Gap. To mitigate these gaps, we introduce PathoTune, a framework designed to efficiently adapt pathological or even visual foundation models to pathology-specific tasks via multi-modal prompt tuning. The proposed framework leverages Task-specific Visual Prompts and Task-specific Textual Prompts to identify task-relevant features, along with Instance-specific Visual Prompts for encoding single pathological image features. Results across multiple datasets at both patch-level and WSI-level demonstrate its superior performance over single-modality prompt tuning approaches. Significantly, PathoTune facilitates the direct adaptation of natural visual foundation models to pathological tasks, drastically outperforming pathological foundation models with simple linear probing. The code is available at https://github.com/openmedlab/PathoDuet.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16497
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PathoTune: Adapting Visual Foundation Model to Pathological Specialists
Lu, Jiaxuan
Yan, Fang
Zhang, Xiaofan
Gao, Yue
Zhang, Shaoting
Computer Vision and Pattern Recognition
Machine Learning
As natural image understanding moves towards the pretrain-finetune era, research in pathology imaging is concurrently evolving. Despite the predominant focus on pretraining pathological foundation models, how to adapt foundation models to downstream tasks is little explored. For downstream adaptation, we propose the existence of two domain gaps, i.e., the Foundation-Task Gap and the Task-Instance Gap. To mitigate these gaps, we introduce PathoTune, a framework designed to efficiently adapt pathological or even visual foundation models to pathology-specific tasks via multi-modal prompt tuning. The proposed framework leverages Task-specific Visual Prompts and Task-specific Textual Prompts to identify task-relevant features, along with Instance-specific Visual Prompts for encoding single pathological image features. Results across multiple datasets at both patch-level and WSI-level demonstrate its superior performance over single-modality prompt tuning approaches. Significantly, PathoTune facilitates the direct adaptation of natural visual foundation models to pathological tasks, drastically outperforming pathological foundation models with simple linear probing. The code is available at https://github.com/openmedlab/PathoDuet.
title PathoTune: Adapting Visual Foundation Model to Pathological Specialists
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2403.16497