Understanding the Transfer Limits of Vision Foundation Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Shiqi, Wang, Yipei, Thorley, Natasha, Ng, Alexander, Saeed, Shaheer, Emberton, Mark, Punwani, Shonit, Kasivisvanathan, Veeru, Barratt, Dean, Alexander, Daniel, Hu, Yipeng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911392214286336
author Huang, Shiqi
Wang, Yipei
Thorley, Natasha
Ng, Alexander
Saeed, Shaheer
Emberton, Mark
Punwani, Shonit
Kasivisvanathan, Veeru
Barratt, Dean
Alexander, Daniel
Hu, Yipeng
author_facet Huang, Shiqi
Wang, Yipei
Thorley, Natasha
Ng, Alexander
Saeed, Shaheer
Emberton, Mark
Punwani, Shonit
Kasivisvanathan, Veeru
Barratt, Dean
Alexander, Daniel
Hu, Yipeng
contents Foundation models leverage large-scale pretraining to capture extensive knowledge, demonstrating generalization in a wide range of language tasks. By comparison, vision foundation models (VFMs) often exhibit uneven improvements across downstream tasks, despite substantial computational investment. We postulate that this limitation arises from a mismatch between pretraining objectives and the demands of downstream vision-and-imaging tasks. Pretraining strategies like masked image reconstruction or contrastive learning shape representations for tasks such as recovery of generic visual patterns or global semantic structures, which may not align with the task-specific requirements of downstream applications including segmentation, classification, or image synthesis. To investigate this in a concrete real-world clinical area, we assess two VFMs, a reconstruction-focused MAE-based model (ProFound) and a contrastive-learning-based model (ProViCNet), on five prostate multiparametric MR imaging tasks, examining how such task alignment influences transfer performance, i.e., from pretraining to fine-tuning. Our findings indicate that better alignment between pretraining and downstream tasks, measured by simple divergence metrics such as maximum-mean-discrepancy (MMD) between the same features before and after fine-tuning, correlates with greater performance improvements and faster convergence, emphasizing the importance of designing and analyzing pretraining objectives with downstream applicability in mind.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15888
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Understanding the Transfer Limits of Vision Foundation Models
Huang, Shiqi
Wang, Yipei
Thorley, Natasha
Ng, Alexander
Saeed, Shaheer
Emberton, Mark
Punwani, Shonit
Kasivisvanathan, Veeru
Barratt, Dean
Alexander, Daniel
Hu, Yipeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Foundation models leverage large-scale pretraining to capture extensive knowledge, demonstrating generalization in a wide range of language tasks. By comparison, vision foundation models (VFMs) often exhibit uneven improvements across downstream tasks, despite substantial computational investment. We postulate that this limitation arises from a mismatch between pretraining objectives and the demands of downstream vision-and-imaging tasks. Pretraining strategies like masked image reconstruction or contrastive learning shape representations for tasks such as recovery of generic visual patterns or global semantic structures, which may not align with the task-specific requirements of downstream applications including segmentation, classification, or image synthesis. To investigate this in a concrete real-world clinical area, we assess two VFMs, a reconstruction-focused MAE-based model (ProFound) and a contrastive-learning-based model (ProViCNet), on five prostate multiparametric MR imaging tasks, examining how such task alignment influences transfer performance, i.e., from pretraining to fine-tuning. Our findings indicate that better alignment between pretraining and downstream tasks, measured by simple divergence metrics such as maximum-mean-discrepancy (MMD) between the same features before and after fine-tuning, correlates with greater performance improvements and faster convergence, emphasizing the importance of designing and analyzing pretraining objectives with downstream applicability in mind.
title Understanding the Transfer Limits of Vision Foundation Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.15888