EXACT: an explainable anomaly-aware vision foundation model for analysis of 3D chest CT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Xuguang, Liu, Mingxuan, Song, Tongxi, Chen, Yifei, Yang, Hongjia, Anmahapong, Kasidit, Li, Zihan, Zhou, Ying, Tian, Qiyuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910168705400832
author Bai, Xuguang
Liu, Mingxuan
Song, Tongxi
Chen, Yifei
Yang, Hongjia
Anmahapong, Kasidit
Li, Zihan
Zhou, Ying
Tian, Qiyuan
author_facet Bai, Xuguang
Liu, Mingxuan
Song, Tongxi
Chen, Yifei
Yang, Hongjia
Anmahapong, Kasidit
Li, Zihan
Zhou, Ying
Tian, Qiyuan
contents Chest computed tomography (CT) is central to the detection and management of thoracic disease, yet the growing scale and complexity of volumetric imaging increasingly exceed what can be addressed by scan-level prediction alone. Clinically useful AI for CT must not only recognize disease across the whole volume, but also localize abnormalities and provide interpretable visual evidence. Existing vision-language foundation models typically compress scans and reports into global image-text representations, limiting their ability to preserve spatial evidence and support clinically meaningful interpretation. Here we developed EXACT, an explainable anomaly-aware foundation model for three-dimensional chest CT that learns spatially resolved representations from paired clinical scans and radiology reports. EXACT was pre-trained on 25,692 CT-reports pairs using anatomy-aware weak supervision, jointly learning organ segmentation and multi-instance anomaly localization without manual voxel-level annotations. The resulting organ-specific anomaly-aware maps assign each voxel a disease-specific anomaly score confined to its corresponding anatomy, jointly encoding lesion extent and organ-level context. In retrospective multinational and multi-center evaluations, EXACT showed broad and consistent improvements across clinically relevant CT tasks, spanning multi-disease diagnosis, zero-shot anomaly localization, downstream adaptation, and visually grounded report generation, outperforming existing three-dimensional medical foundation models. By transforming routine clinical CT scans and free-text reports into explainable voxel-level representations, EXACT establishes a scalable paradigm for trustworthy volumetric medical AI.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24146
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EXACT: an explainable anomaly-aware vision foundation model for analysis of 3D chest CT
Bai, Xuguang
Liu, Mingxuan
Song, Tongxi
Chen, Yifei
Yang, Hongjia
Anmahapong, Kasidit
Li, Zihan
Zhou, Ying
Tian, Qiyuan
Computer Vision and Pattern Recognition
Chest computed tomography (CT) is central to the detection and management of thoracic disease, yet the growing scale and complexity of volumetric imaging increasingly exceed what can be addressed by scan-level prediction alone. Clinically useful AI for CT must not only recognize disease across the whole volume, but also localize abnormalities and provide interpretable visual evidence. Existing vision-language foundation models typically compress scans and reports into global image-text representations, limiting their ability to preserve spatial evidence and support clinically meaningful interpretation. Here we developed EXACT, an explainable anomaly-aware foundation model for three-dimensional chest CT that learns spatially resolved representations from paired clinical scans and radiology reports. EXACT was pre-trained on 25,692 CT-reports pairs using anatomy-aware weak supervision, jointly learning organ segmentation and multi-instance anomaly localization without manual voxel-level annotations. The resulting organ-specific anomaly-aware maps assign each voxel a disease-specific anomaly score confined to its corresponding anatomy, jointly encoding lesion extent and organ-level context. In retrospective multinational and multi-center evaluations, EXACT showed broad and consistent improvements across clinically relevant CT tasks, spanning multi-disease diagnosis, zero-shot anomaly localization, downstream adaptation, and visually grounded report generation, outperforming existing three-dimensional medical foundation models. By transforming routine clinical CT scans and free-text reports into explainable voxel-level representations, EXACT establishes a scalable paradigm for trustworthy volumetric medical AI.
title EXACT: an explainable anomaly-aware vision foundation model for analysis of 3D chest CT
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.24146