Disease-informed Adaptation of Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiajin, Wang, Ge, Kalra, Mannudeep K., Yan, Pingkun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917674345299968
author Zhang, Jiajin
Wang, Ge
Kalra, Mannudeep K.
Yan, Pingkun
author_facet Zhang, Jiajin
Wang, Ge
Kalra, Mannudeep K.
Yan, Pingkun
contents In medical image analysis, the expertise scarcity and the high cost of data annotation limits the development of large artificial intelligence models. This paper investigates the potential of transfer learning with pre-trained vision-language models (VLMs) in this domain. Currently, VLMs still struggle to transfer to the underrepresented diseases with minimal presence and new diseases entirely absent from the pretraining dataset. We argue that effective adaptation of VLMs hinges on the nuanced representation learning of disease concepts. By capitalizing on the joint visual-linguistic capabilities of VLMs, we introduce disease-informed contextual prompting in a novel disease prototype learning framework. This approach enables VLMs to grasp the concepts of new disease effectively and efficiently, even with limited data. Extensive experiments across multiple image modalities showcase notable enhancements in performance compared to existing techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15728
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Disease-informed Adaptation of Vision-Language Models
Zhang, Jiajin
Wang, Ge
Kalra, Mannudeep K.
Yan, Pingkun
Computer Vision and Pattern Recognition
In medical image analysis, the expertise scarcity and the high cost of data annotation limits the development of large artificial intelligence models. This paper investigates the potential of transfer learning with pre-trained vision-language models (VLMs) in this domain. Currently, VLMs still struggle to transfer to the underrepresented diseases with minimal presence and new diseases entirely absent from the pretraining dataset. We argue that effective adaptation of VLMs hinges on the nuanced representation learning of disease concepts. By capitalizing on the joint visual-linguistic capabilities of VLMs, we introduce disease-informed contextual prompting in a novel disease prototype learning framework. This approach enables VLMs to grasp the concepts of new disease effectively and efficiently, even with limited data. Extensive experiments across multiple image modalities showcase notable enhancements in performance compared to existing techniques.
title Disease-informed Adaptation of Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.15728