Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bawazir, Ameera, Wu, Kebin, Li, Wenbin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929601838579712
author Bawazir, Ameera
Wu, Kebin
Li, Wenbin
author_facet Bawazir, Ameera
Wu, Kebin
Li, Wenbin
contents Recent advancements in vision-language pre-training via contrastive learning have significantly improved performance across computer vision tasks. However, in the medical domain, obtaining multimodal data is often costly and challenging due to privacy, sensitivity, and annotation complexity. To mitigate data scarcity while boosting model performance, we introduce \textbf{Uni-Mlip}, a unified self-supervision framework specifically designed to enhance medical vision-language pre-training. Uni-Mlip seamlessly integrates cross-modality, uni-modality, and fused-modality self-supervision techniques at the data-level and the feature-level. Additionally, Uni-Mlip tailors uni-modal image self-supervision to accommodate the unique characteristics of medical images. Our experiments across datasets of varying scales demonstrate that Uni-Mlip significantly surpasses current state-of-the-art methods in three key downstream tasks: image-text retrieval, image classification, and visual question answering (VQA).
format Preprint
id arxiv_https___arxiv_org_abs_2411_15207
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
Bawazir, Ameera
Wu, Kebin
Li, Wenbin
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Recent advancements in vision-language pre-training via contrastive learning have significantly improved performance across computer vision tasks. However, in the medical domain, obtaining multimodal data is often costly and challenging due to privacy, sensitivity, and annotation complexity. To mitigate data scarcity while boosting model performance, we introduce \textbf{Uni-Mlip}, a unified self-supervision framework specifically designed to enhance medical vision-language pre-training. Uni-Mlip seamlessly integrates cross-modality, uni-modality, and fused-modality self-supervision techniques at the data-level and the feature-level. Additionally, Uni-Mlip tailors uni-modal image self-supervision to accommodate the unique characteristics of medical images. Our experiments across datasets of varying scales demonstrate that Uni-Mlip significantly surpasses current state-of-the-art methods in three key downstream tasks: image-text retrieval, image classification, and visual question answering (VQA).
title Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.15207