XGeM: A Multi-Prompt Foundation Model for Multimodal Medical Data Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Molino, Daniele, Di Feola, Francesco, Faiella, Eliodoro, Fazzini, Deborah, Santucci, Domiziana, Shen, Linlin, Guarrasi, Valerio, Soda, Paolo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908832615104512
author Molino, Daniele
Di Feola, Francesco
Faiella, Eliodoro
Fazzini, Deborah
Santucci, Domiziana
Shen, Linlin
Guarrasi, Valerio
Soda, Paolo
author_facet Molino, Daniele
Di Feola, Francesco
Faiella, Eliodoro
Fazzini, Deborah
Santucci, Domiziana
Shen, Linlin
Guarrasi, Valerio
Soda, Paolo
contents The adoption of Artificial Intelligence in medical imaging holds great promise, yet it remains hindered by challenges such as data scarcity, privacy concerns, and the need for robust multimodal integration. While recent advances in generative modeling have enabled high-quality synthetic data generation, existing approaches are often limited to unimodal, unidirectional synthesis and therefore lack the ability to jointly synthesize multiple modalities while preserving clinical consistency. To address this challenge, we introduce XGeM, a 6.77-billion-parameter multimodal generative model designed to support flexible, any-to-any synthesis between medical data modalities. XGeM constructs a shared latent space via contrastive learning and introduces a novel Multi-Prompt Training strategy, enabling conditioning on arbitrary subsets of input modalities. This design allows the model to adapt to heterogeneous clinical inputs and generate multiple outputs jointly, preserving both semantic and structural coherence. We extensively validate XGeM: first we benchmark it against five competitors on the MIMIC-CXR dataset, a state-of-the-art dataset for multi-view Chest X-ray and radiological report generation. Secondly, we perform a Visual Turing Test with expert radiologists to assess the realism and clinical relevance of the generated data, ensuring alignment with real-world scenarios. Finally, we show how XGeM can support key medical data challenges such as anonymization, class imbalance, and data scarcity, underscoring its utility as a foundation model for medical data synthesis. Project page is at https://cosbidev.github.io/XGeM/.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04614
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle XGeM: A Multi-Prompt Foundation Model for Multimodal Medical Data Generation
Molino, Daniele
Di Feola, Francesco
Faiella, Eliodoro
Fazzini, Deborah
Santucci, Domiziana
Shen, Linlin
Guarrasi, Valerio
Soda, Paolo
Artificial Intelligence
Machine Learning
The adoption of Artificial Intelligence in medical imaging holds great promise, yet it remains hindered by challenges such as data scarcity, privacy concerns, and the need for robust multimodal integration. While recent advances in generative modeling have enabled high-quality synthetic data generation, existing approaches are often limited to unimodal, unidirectional synthesis and therefore lack the ability to jointly synthesize multiple modalities while preserving clinical consistency. To address this challenge, we introduce XGeM, a 6.77-billion-parameter multimodal generative model designed to support flexible, any-to-any synthesis between medical data modalities. XGeM constructs a shared latent space via contrastive learning and introduces a novel Multi-Prompt Training strategy, enabling conditioning on arbitrary subsets of input modalities. This design allows the model to adapt to heterogeneous clinical inputs and generate multiple outputs jointly, preserving both semantic and structural coherence. We extensively validate XGeM: first we benchmark it against five competitors on the MIMIC-CXR dataset, a state-of-the-art dataset for multi-view Chest X-ray and radiological report generation. Secondly, we perform a Visual Turing Test with expert radiologists to assess the realism and clinical relevance of the generated data, ensuring alignment with real-world scenarios. Finally, we show how XGeM can support key medical data challenges such as anonymization, class imbalance, and data scarcity, underscoring its utility as a foundation model for medical data synthesis. Project page is at https://cosbidev.github.io/XGeM/.
title XGeM: A Multi-Prompt Foundation Model for Multimodal Medical Data Generation
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2501.04614