Distribution-Based Masked Medical Vision-Language Model Using Structured Reports

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gowda, Shreyank N, Zhang, Ruichi, Gu, Xiao, Weng, Ying, Yang, Lu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909710783873024
author Gowda, Shreyank N
Zhang, Ruichi
Gu, Xiao
Weng, Ying
Yang, Lu
author_facet Gowda, Shreyank N
Zhang, Ruichi
Gu, Xiao
Weng, Ying
Yang, Lu
contents Medical image-language pre-training aims to align medical images with clinically relevant text to improve model performance on various downstream tasks. However, existing models often struggle with the variability and ambiguity inherent in medical data, limiting their ability to capture nuanced clinical information and uncertainty. This work introduces an uncertainty-aware medical image-text pre-training model that enhances generalization capabilities in medical image analysis. Building on previous methods and focusing on Chest X-Rays, our approach utilizes structured text reports generated by a large language model (LLM) to augment image data with clinically relevant context. These reports begin with a definition of the disease, followed by the `appearance' section to highlight critical regions of interest, and finally `observations' and `verdicts' that ground model predictions in clinical semantics. By modeling both inter- and intra-modal uncertainty, our framework captures the inherent ambiguity in medical images and text, yielding improved representations and performance on downstream tasks. Our model demonstrates significant advances in medical image-text pre-training, obtaining state-of-the-art performance on multiple downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21794
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Distribution-Based Masked Medical Vision-Language Model Using Structured Reports
Gowda, Shreyank N
Zhang, Ruichi
Gu, Xiao
Weng, Ying
Yang, Lu
Computer Vision and Pattern Recognition
Medical image-language pre-training aims to align medical images with clinically relevant text to improve model performance on various downstream tasks. However, existing models often struggle with the variability and ambiguity inherent in medical data, limiting their ability to capture nuanced clinical information and uncertainty. This work introduces an uncertainty-aware medical image-text pre-training model that enhances generalization capabilities in medical image analysis. Building on previous methods and focusing on Chest X-Rays, our approach utilizes structured text reports generated by a large language model (LLM) to augment image data with clinically relevant context. These reports begin with a definition of the disease, followed by the `appearance' section to highlight critical regions of interest, and finally `observations' and `verdicts' that ground model predictions in clinical semantics. By modeling both inter- and intra-modal uncertainty, our framework captures the inherent ambiguity in medical images and text, yielding improved representations and performance on downstream tasks. Our model demonstrates significant advances in medical image-text pre-training, obtaining state-of-the-art performance on multiple downstream tasks.
title Distribution-Based Masked Medical Vision-Language Model Using Structured Reports
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.21794