A Generative Foundation Model for Chest Radiography

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ji, Yuanfeng, Lin, Dan, Wang, Xiyue, Zhang, Lu, Zhou, Wenhui, Ge, Chongjian, Chu, Ruihang, Yang, Xiaoli, Zhao, Junhan, Chen, Junsong, Luo, Xiangde, Yang, Sen, Fang, Jin, Luo, Ping, Li, Ruijiang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915478654418944
author Ji, Yuanfeng
Lin, Dan
Wang, Xiyue
Zhang, Lu
Zhou, Wenhui
Ge, Chongjian
Chu, Ruihang
Yang, Xiaoli
Zhao, Junhan
Chen, Junsong
Luo, Xiangde
Yang, Sen
Fang, Jin
Luo, Ping
Li, Ruijiang
author_facet Ji, Yuanfeng
Lin, Dan
Wang, Xiyue
Zhang, Lu
Zhou, Wenhui
Ge, Chongjian
Chu, Ruihang
Yang, Xiaoli
Zhao, Junhan
Chen, Junsong
Luo, Xiangde
Yang, Sen
Fang, Jin
Luo, Ping
Li, Ruijiang
contents The scarcity of well-annotated diverse medical images is a major hurdle for developing reliable AI models in healthcare. Substantial technical advances have been made in generative foundation models for natural images. Here we develop `ChexGen', a generative vision-language foundation model that introduces a unified framework for text-, mask-, and bounding box-guided synthesis of chest radiographs. Built upon the latent diffusion transformer architecture, ChexGen was pretrained on the largest curated chest X-ray dataset to date, consisting of 960,000 radiograph-report pairs. ChexGen achieves accurate synthesis of radiographs through expert evaluations and quantitative metrics. We demonstrate the utility of ChexGen for training data augmentation and supervised pretraining, which led to performance improvements across disease classification, detection, and segmentation tasks using a small fraction of training data. Further, our model enables the creation of diverse patient cohorts that enhance model fairness by detecting and mitigating demographic biases. Our study supports the transformative role of generative foundation models in building more accurate, data-efficient, and equitable medical AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03903
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Generative Foundation Model for Chest Radiography
Ji, Yuanfeng
Lin, Dan
Wang, Xiyue
Zhang, Lu
Zhou, Wenhui
Ge, Chongjian
Chu, Ruihang
Yang, Xiaoli
Zhao, Junhan
Chen, Junsong
Luo, Xiangde
Yang, Sen
Fang, Jin
Luo, Ping
Li, Ruijiang
Computer Vision and Pattern Recognition
The scarcity of well-annotated diverse medical images is a major hurdle for developing reliable AI models in healthcare. Substantial technical advances have been made in generative foundation models for natural images. Here we develop `ChexGen', a generative vision-language foundation model that introduces a unified framework for text-, mask-, and bounding box-guided synthesis of chest radiographs. Built upon the latent diffusion transformer architecture, ChexGen was pretrained on the largest curated chest X-ray dataset to date, consisting of 960,000 radiograph-report pairs. ChexGen achieves accurate synthesis of radiographs through expert evaluations and quantitative metrics. We demonstrate the utility of ChexGen for training data augmentation and supervised pretraining, which led to performance improvements across disease classification, detection, and segmentation tasks using a small fraction of training data. Further, our model enables the creation of diverse patient cohorts that enhance model fairness by detecting and mitigating demographic biases. Our study supports the transformative role of generative foundation models in building more accurate, data-efficient, and equitable medical AI systems.
title A Generative Foundation Model for Chest Radiography
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.03903