Xray2Xray: World Model from Chest X-rays with Volumetric Context

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zefan, Song, Xinrui, Xu, Xuanang, Shi, Yongyi, Wang, Ge, Kalra, Mannudeep K., Yan, Pingkun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909658613022720
author Yang, Zefan
Song, Xinrui
Xu, Xuanang
Shi, Yongyi
Wang, Ge
Kalra, Mannudeep K.
Yan, Pingkun
author_facet Yang, Zefan
Song, Xinrui
Xu, Xuanang
Shi, Yongyi
Wang, Ge
Kalra, Mannudeep K.
Yan, Pingkun
contents Chest X-rays (CXRs) are the most widely used medical imaging modality and play a pivotal role in diagnosing diseases. However, as 2D projection images, CXRs are limited by structural superposition, which constrains their effectiveness in precise disease diagnosis and risk prediction. To address the limitations of 2D CXRs, this study introduces Xray2Xray, a novel World Model that learns latent representations encoding 3D structural information from chest X-rays. Xray2Xray captures the latent representations of the chest volume by modeling the transition dynamics of X-ray projections across different angular positions with a vision model and a transition model. We employed the latent representations of Xray2Xray for downstream risk prediction and disease diagnosis tasks. Experimental results showed that Xray2Xray outperformed both supervised methods and self-supervised pretraining methods for cardiovascular disease risk estimation and achieved competitive performance in classifying five pathologies in CXRs. We also assessed the quality of Xray2Xray's latent representations through synthesis tasks and demonstrated that the latent representations can be used to reconstruct volumetric context.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Xray2Xray: World Model from Chest X-rays with Volumetric Context
Yang, Zefan
Song, Xinrui
Xu, Xuanang
Shi, Yongyi
Wang, Ge
Kalra, Mannudeep K.
Yan, Pingkun
Image and Video Processing
Computer Vision and Pattern Recognition
Chest X-rays (CXRs) are the most widely used medical imaging modality and play a pivotal role in diagnosing diseases. However, as 2D projection images, CXRs are limited by structural superposition, which constrains their effectiveness in precise disease diagnosis and risk prediction. To address the limitations of 2D CXRs, this study introduces Xray2Xray, a novel World Model that learns latent representations encoding 3D structural information from chest X-rays. Xray2Xray captures the latent representations of the chest volume by modeling the transition dynamics of X-ray projections across different angular positions with a vision model and a transition model. We employed the latent representations of Xray2Xray for downstream risk prediction and disease diagnosis tasks. Experimental results showed that Xray2Xray outperformed both supervised methods and self-supervised pretraining methods for cardiovascular disease risk estimation and achieved competitive performance in classifying five pathologies in CXRs. We also assessed the quality of Xray2Xray's latent representations through synthesis tasks and demonstrated that the latent representations can be used to reconstruct volumetric context.
title Xray2Xray: World Model from Chest X-rays with Volumetric Context
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.19055