ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Ziqing, Liang, Cheng, Wu, Chaoyi, Zhang, Ya, Wang, Yanfeng, Xie, Weidi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909617827610624
author Fan, Ziqing
Liang, Cheng
Wu, Chaoyi
Zhang, Ya
Wang, Yanfeng
Xie, Weidi
author_facet Fan, Ziqing
Liang, Cheng
Wu, Chaoyi
Zhang, Ya
Wang, Yanfeng
Xie, Weidi
contents Recent advances in reasoning-enhanced large language models (LLMs) and multimodal LLMs (MLLMs) have significantly improved performance in complex tasks, yet medical AI models often overlook the structured reasoning processes inherent in clinical practice. In this work, we present ChestX-Reasoner, a radiology diagnosis MLLM designed to leverage process supervision mined directly from clinical reports, reflecting the step-by-step reasoning followed by radiologists. We construct a large dataset by extracting and refining reasoning chains from routine radiology reports. Our two-stage training framework combines supervised fine-tuning and reinforcement learning guided by process rewards to better align model reasoning with clinical standards. We introduce RadRBench-CXR, a comprehensive benchmark featuring 59K visual question answering samples with 301K clinically validated reasoning steps, and propose RadRScore, a metric evaluating reasoning factuality, completeness, and effectiveness. ChestX-Reasoner outperforms existing medical and general-domain MLLMs in both diagnostic accuracy and reasoning ability, achieving 16%, 5.9%, and 18% improvements in reasoning ability compared to the best medical MLLM, the best general MLLM, and its base model, respectively, as well as 3.3%, 24%, and 27% improvements in outcome accuracy. All resources are open-sourced to facilitate further research in medical reasoning MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2504_20930
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
Fan, Ziqing
Liang, Cheng
Wu, Chaoyi
Zhang, Ya
Wang, Yanfeng
Xie, Weidi
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Recent advances in reasoning-enhanced large language models (LLMs) and multimodal LLMs (MLLMs) have significantly improved performance in complex tasks, yet medical AI models often overlook the structured reasoning processes inherent in clinical practice. In this work, we present ChestX-Reasoner, a radiology diagnosis MLLM designed to leverage process supervision mined directly from clinical reports, reflecting the step-by-step reasoning followed by radiologists. We construct a large dataset by extracting and refining reasoning chains from routine radiology reports. Our two-stage training framework combines supervised fine-tuning and reinforcement learning guided by process rewards to better align model reasoning with clinical standards. We introduce RadRBench-CXR, a comprehensive benchmark featuring 59K visual question answering samples with 301K clinically validated reasoning steps, and propose RadRScore, a metric evaluating reasoning factuality, completeness, and effectiveness. ChestX-Reasoner outperforms existing medical and general-domain MLLMs in both diagnostic accuracy and reasoning ability, achieving 16%, 5.9%, and 18% improvements in reasoning ability compared to the best medical MLLM, the best general MLLM, and its base model, respectively, as well as 3.3%, 24%, and 27% improvements in outcome accuracy. All resources are open-sourced to facilitate further research in medical reasoning MLLMs.
title ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.20930