Deep Reprogramming Distillation for Medical Foundation Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Du, Siyuan, Zhou, Yuhang, Li, Haolin, Yao, Jiangchao, Wang, Haishuai, Lin, Hui, Zhang, Ya, Wang, Yanfeng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917463314137088
author Du, Siyuan
Zhou, Yuhang
Li, Haolin
Yao, Jiangchao
Wang, Haishuai
Lin, Hui
Zhang, Ya
Wang, Yanfeng
author_facet Du, Siyuan
Zhou, Yuhang
Li, Haolin
Yao, Jiangchao
Wang, Haishuai
Lin, Hui
Zhang, Ya
Wang, Yanfeng
contents Medical foundation models pre-trained on large-scale datasets have shown powerful versatile performance. However, when adapting medical foundation models for specific medical scenarios, it remains the inevitable challenge due to the gap induced by the discrepancy between pre-training and downstream tasks, the real-world computation, and speed constraints. Relevant techniques that probably handle this challenge more or less suffer from some intrinsic limitations. For example, knowledge distillation (KD) assumes that teacher and student models share the same task, training strategy, and model structure family, while prevalent parameter-efficient fine-tuning (PEFT) fails to achieve personalized and lightweight deployment. Even the combination of PEFT and KD still struggles to resolve model structures and training strategies inconsistencies between teacher and student models, leading to inefficient knowledge transfer. In this study, we propose a novel framework called Deep Reprogramming Distillation (DRD) to combat the general adaptation challenge. Specifically, DRD introduces the novel reprogramming module that on the one side overcomes the domain and task discrepancy between pretraining and downstream scenarios, and on the other side builds the student-friendly efficient distillation from foundation models to lightweight downstream models. Furthermore, to mitigate variability under different training conditions, we design a centered kernel alignment (CKA) distillation method to promote robust knowledge transfer. Empirical results show that DRD surpasses previous PEFT and KD methods across 18 medical downstream tasks under different foundation models, covering various scenarios including 2D/3D classification and 2D/3D segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04447
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Deep Reprogramming Distillation for Medical Foundation Models
Du, Siyuan
Zhou, Yuhang
Li, Haolin
Yao, Jiangchao
Wang, Haishuai
Lin, Hui
Zhang, Ya
Wang, Yanfeng
Computer Vision and Pattern Recognition
Medical foundation models pre-trained on large-scale datasets have shown powerful versatile performance. However, when adapting medical foundation models for specific medical scenarios, it remains the inevitable challenge due to the gap induced by the discrepancy between pre-training and downstream tasks, the real-world computation, and speed constraints. Relevant techniques that probably handle this challenge more or less suffer from some intrinsic limitations. For example, knowledge distillation (KD) assumes that teacher and student models share the same task, training strategy, and model structure family, while prevalent parameter-efficient fine-tuning (PEFT) fails to achieve personalized and lightweight deployment. Even the combination of PEFT and KD still struggles to resolve model structures and training strategies inconsistencies between teacher and student models, leading to inefficient knowledge transfer. In this study, we propose a novel framework called Deep Reprogramming Distillation (DRD) to combat the general adaptation challenge. Specifically, DRD introduces the novel reprogramming module that on the one side overcomes the domain and task discrepancy between pretraining and downstream scenarios, and on the other side builds the student-friendly efficient distillation from foundation models to lightweight downstream models. Furthermore, to mitigate variability under different training conditions, we design a centered kernel alignment (CKA) distillation method to promote robust knowledge transfer. Empirical results show that DRD surpasses previous PEFT and KD methods across 18 medical downstream tasks under different foundation models, covering various scenarios including 2D/3D classification and 2D/3D segmentation.
title Deep Reprogramming Distillation for Medical Foundation Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.04447