PuzzleTuning: Explicitly Bridge Pathological and Natural Image with Puzzles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tianyi, Lyu, Shangqing, Lei, Yanli, Chen, Sicheng, Ying, Nan, He, Yufang, Zhao, Yu, Feng, Yunlu, Lee, Hwee Kuan, Zhang, Guanglei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909177803177984
author Zhang, Tianyi
Lyu, Shangqing
Lei, Yanli
Chen, Sicheng
Ying, Nan
He, Yufang
Zhao, Yu
Feng, Yunlu
Lee, Hwee Kuan
Zhang, Guanglei
author_facet Zhang, Tianyi
Lyu, Shangqing
Lei, Yanli
Chen, Sicheng
Ying, Nan
He, Yufang
Zhao, Yu
Feng, Yunlu
Lee, Hwee Kuan
Zhang, Guanglei
contents Pathological image analysis is a crucial field in computer vision. Due to the annotation scarcity in the pathological field, pre-training with self-supervised learning (SSL) is widely applied to learn on unlabeled images. However, the current SSL-based pathological pre-training: (1) does not explicitly explore the essential focuses of the pathological field, and (2) does not effectively bridge with and thus take advantage of the knowledge from natural images. To explicitly address them, we propose our large-scale PuzzleTuning framework, containing the following innovations. Firstly, we define three task focuses that can effectively bridge knowledge of pathological and natural domain: appearance consistency, spatial consistency, and restoration understanding. Secondly, we devise a novel multiple puzzle restoring task, which explicitly pre-trains the model regarding these focuses. Thirdly, we introduce an explicit prompt-tuning process to incrementally integrate the domain-specific knowledge. It builds a bridge to align the large domain gap between natural and pathological images. Additionally, a curriculum-learning training strategy is designed to regulate task difficulty, making the model adaptive to the puzzle restoring complexity. Experimental results show that our PuzzleTuning framework outperforms the previous state-of-the-art methods in various downstream tasks on multiple datasets. The code, demo, and pre-trained weights are available at https://github.com/sagizty/PuzzleTuning.
format Preprint
id arxiv_https___arxiv_org_abs_2311_06712
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle PuzzleTuning: Explicitly Bridge Pathological and Natural Image with Puzzles
Zhang, Tianyi
Lyu, Shangqing
Lei, Yanli
Chen, Sicheng
Ying, Nan
He, Yufang
Zhao, Yu
Feng, Yunlu
Lee, Hwee Kuan
Zhang, Guanglei
Image and Video Processing
Pathological image analysis is a crucial field in computer vision. Due to the annotation scarcity in the pathological field, pre-training with self-supervised learning (SSL) is widely applied to learn on unlabeled images. However, the current SSL-based pathological pre-training: (1) does not explicitly explore the essential focuses of the pathological field, and (2) does not effectively bridge with and thus take advantage of the knowledge from natural images. To explicitly address them, we propose our large-scale PuzzleTuning framework, containing the following innovations. Firstly, we define three task focuses that can effectively bridge knowledge of pathological and natural domain: appearance consistency, spatial consistency, and restoration understanding. Secondly, we devise a novel multiple puzzle restoring task, which explicitly pre-trains the model regarding these focuses. Thirdly, we introduce an explicit prompt-tuning process to incrementally integrate the domain-specific knowledge. It builds a bridge to align the large domain gap between natural and pathological images. Additionally, a curriculum-learning training strategy is designed to regulate task difficulty, making the model adaptive to the puzzle restoring complexity. Experimental results show that our PuzzleTuning framework outperforms the previous state-of-the-art methods in various downstream tasks on multiple datasets. The code, demo, and pre-trained weights are available at https://github.com/sagizty/PuzzleTuning.
title PuzzleTuning: Explicitly Bridge Pathological and Natural Image with Puzzles
topic Image and Video Processing
url https://arxiv.org/abs/2311.06712