Evaluating Strategies for Synthesizing Clinical Notes for Medical Multimodal AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Marini, Niccolo, Liang, Zhaohui, Rajaraman, Sivaramakrishnan, Xue, Zhiyun, Antani, Sameer
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908677605163008
author Marini, Niccolo
Liang, Zhaohui
Rajaraman, Sivaramakrishnan
Xue, Zhiyun
Antani, Sameer
author_facet Marini, Niccolo
Liang, Zhaohui
Rajaraman, Sivaramakrishnan
Xue, Zhiyun
Antani, Sameer
contents Multimodal (MM) learning is emerging as a promising paradigm in biomedical artificial intelligence (AI) applications, integrating complementary modality, which highlight different aspects of patient health. The scarcity of large heterogeneous biomedical MM data has restrained the development of robust models for medical AI applications. In the dermatology domain, for instance, skin lesion datasets typically include only images linked to minimal metadata describing the condition, thereby limiting the benefits of MM data integration for reliable and generalizable predictions. Recent advances in Large Language Models (LLMs) enable the synthesis of textual description of image findings, potentially allowing the combination of image and text representations. However, LLMs are not specifically trained for use in the medical domain, and their naive inclusion has raised concerns about the risk of hallucinations in clinically relevant contexts. This work investigates strategies for generating synthetic textual clinical notes, in terms of prompt design and medical metadata inclusion, and evaluates their impact on MM architectures toward enhancing performance in classification and cross-modal retrieval tasks. Experiments across several heterogeneous dermatology datasets demonstrate that synthetic clinical notes not only enhance classification performance, particularly under domain shift, but also unlock cross-modal retrieval capabilities, a downstream task that is not explicitly optimized during training.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Strategies for Synthesizing Clinical Notes for Medical Multimodal AI
Marini, Niccolo
Liang, Zhaohui
Rajaraman, Sivaramakrishnan
Xue, Zhiyun
Antani, Sameer
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimodal (MM) learning is emerging as a promising paradigm in biomedical artificial intelligence (AI) applications, integrating complementary modality, which highlight different aspects of patient health. The scarcity of large heterogeneous biomedical MM data has restrained the development of robust models for medical AI applications. In the dermatology domain, for instance, skin lesion datasets typically include only images linked to minimal metadata describing the condition, thereby limiting the benefits of MM data integration for reliable and generalizable predictions. Recent advances in Large Language Models (LLMs) enable the synthesis of textual description of image findings, potentially allowing the combination of image and text representations. However, LLMs are not specifically trained for use in the medical domain, and their naive inclusion has raised concerns about the risk of hallucinations in clinically relevant contexts. This work investigates strategies for generating synthetic textual clinical notes, in terms of prompt design and medical metadata inclusion, and evaluates their impact on MM architectures toward enhancing performance in classification and cross-modal retrieval tasks. Experiments across several heterogeneous dermatology datasets demonstrate that synthetic clinical notes not only enhance classification performance, particularly under domain shift, but also unlock cross-modal retrieval capabilities, a downstream task that is not explicitly optimized during training.
title Evaluating Strategies for Synthesizing Clinical Notes for Medical Multimodal AI
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.21827