HySparK: Hybrid Sparse Masking for Large Scale Medical Image Pre-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Fenghe, Xu, Ronghao, Yao, Qingsong, Fu, Xueming, Quan, Quan, Zhu, Heqin, Liu, Zaiyi, Zhou, S. Kevin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916353182531584
author Tang, Fenghe
Xu, Ronghao
Yao, Qingsong
Fu, Xueming
Quan, Quan
Zhu, Heqin
Liu, Zaiyi
Zhou, S. Kevin
author_facet Tang, Fenghe
Xu, Ronghao
Yao, Qingsong
Fu, Xueming
Quan, Quan
Zhu, Heqin
Liu, Zaiyi
Zhou, S. Kevin
contents The generative self-supervised learning strategy exhibits remarkable learning representational capabilities. However, there is limited attention to end-to-end pre-training methods based on a hybrid architecture of CNN and Transformer, which can learn strong local and global representations simultaneously. To address this issue, we propose a generative pre-training strategy called Hybrid Sparse masKing (HySparK) based on masked image modeling and apply it to large-scale pre-training on medical images. First, we perform a bottom-up 3D hybrid masking strategy on the encoder to keep consistency masking. Then we utilize sparse convolution for the top CNNs and encode unmasked patches for the bottom vision Transformers. Second, we employ a simple hierarchical decoder with skip-connections to achieve dense multi-scale feature reconstruction. Third, we implement our pre-training method on a collection of multiple large-scale 3D medical imaging datasets. Extensive experiments indicate that our proposed pre-training strategy demonstrates robust transfer-ability in supervised downstream tasks and sheds light on HySparK's promising prospects. The code is available at https://github.com/FengheTan9/HySparK
format Preprint
id arxiv_https___arxiv_org_abs_2408_05815
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HySparK: Hybrid Sparse Masking for Large Scale Medical Image Pre-Training
Tang, Fenghe
Xu, Ronghao
Yao, Qingsong
Fu, Xueming
Quan, Quan
Zhu, Heqin
Liu, Zaiyi
Zhou, S. Kevin
Computer Vision and Pattern Recognition
I.4.10; I.4.6
The generative self-supervised learning strategy exhibits remarkable learning representational capabilities. However, there is limited attention to end-to-end pre-training methods based on a hybrid architecture of CNN and Transformer, which can learn strong local and global representations simultaneously. To address this issue, we propose a generative pre-training strategy called Hybrid Sparse masKing (HySparK) based on masked image modeling and apply it to large-scale pre-training on medical images. First, we perform a bottom-up 3D hybrid masking strategy on the encoder to keep consistency masking. Then we utilize sparse convolution for the top CNNs and encode unmasked patches for the bottom vision Transformers. Second, we employ a simple hierarchical decoder with skip-connections to achieve dense multi-scale feature reconstruction. Third, we implement our pre-training method on a collection of multiple large-scale 3D medical imaging datasets. Extensive experiments indicate that our proposed pre-training strategy demonstrates robust transfer-ability in supervised downstream tasks and sheds light on HySparK's promising prospects. The code is available at https://github.com/FengheTan9/HySparK
title HySparK: Hybrid Sparse Masking for Large Scale Medical Image Pre-Training
topic Computer Vision and Pattern Recognition
I.4.10; I.4.6
url https://arxiv.org/abs/2408.05815