Toward Clinically Explainable AI for Medical Diagnosis: A Foundation Model with Human-Compatible Reasoning via Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lin, Qika, Zhu, Yifan, Pu, Bin, Huang, Ling, Luo, Haoran, Ma, Jingying, Wu, Feng, He, Kai, Xu, Jiaxing, Peng, Zhen, Zhao, Tianzhe, Xu, Fangzhi, Zhang, Jian, Ou, Zhonghong, Cambria, Erik, Mishra, Swapnil, Feng, Mengling
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911477451980800
author Lin, Qika
Zhu, Yifan
Pu, Bin
Huang, Ling
Luo, Haoran
Ma, Jingying
Wu, Feng
He, Kai
Xu, Jiaxing
Peng, Zhen
Zhao, Tianzhe
Xu, Fangzhi
Zhang, Jian
Ou, Zhonghong
Cambria, Erik
Mishra, Swapnil
Feng, Mengling
author_facet Lin, Qika
Zhu, Yifan
Pu, Bin
Huang, Ling
Luo, Haoran
Ma, Jingying
Wu, Feng
He, Kai
Xu, Jiaxing
Peng, Zhen
Zhao, Tianzhe
Xu, Fangzhi
Zhang, Jian
Ou, Zhonghong
Cambria, Erik
Mishra, Swapnil
Feng, Mengling
contents The clinical adoption of artificial intelligence (AI) in medical diagnostics is critically hampered by its black-box nature, which prevents clinicians from verifying the rationale behind automated decisions. To overcome this fundamental barrier, we introduce DeepMedix-R1, a foundation model (FM) for chest X-ray (CXR) interpretation that generates not only accurate diagnoses but also a transparent, step-by-step reasoning process grounded in specific visual evidence. Our methodology employs a sequential training strategy, beginning with instruction fine-tuning, followed by a cold-start phase to elicit reasoning capabilities. Critically, we then implement reinforcement learning with grounded rewards to meticulously refine the model, aligning both its diagnostic outputs and its reasoning pathways with clinical plausibility. Quantitative assessments show that DeepMedix-R1 substantially outperforms advanced FMs, achieving improvements in report generation and visual question answering tasks. We also introduce Report Arena, a novel LLM-based benchmark that ranks DeepMedix-R1 first among competing models for output quality. Most significantly, a formal review by clinical experts reveals a profound preference for DeepMedix-R1's generated reasoning over the broadly adopted Qwen2.5-VL-7B model, confirming its superior interpretability and clinical utility.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03906
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toward Clinically Explainable AI for Medical Diagnosis: A Foundation Model with Human-Compatible Reasoning via Reinforcement Learning
Lin, Qika
Zhu, Yifan
Pu, Bin
Huang, Ling
Luo, Haoran
Ma, Jingying
Wu, Feng
He, Kai
Xu, Jiaxing
Peng, Zhen
Zhao, Tianzhe
Xu, Fangzhi
Zhang, Jian
Ou, Zhonghong
Cambria, Erik
Mishra, Swapnil
Feng, Mengling
Artificial Intelligence
The clinical adoption of artificial intelligence (AI) in medical diagnostics is critically hampered by its black-box nature, which prevents clinicians from verifying the rationale behind automated decisions. To overcome this fundamental barrier, we introduce DeepMedix-R1, a foundation model (FM) for chest X-ray (CXR) interpretation that generates not only accurate diagnoses but also a transparent, step-by-step reasoning process grounded in specific visual evidence. Our methodology employs a sequential training strategy, beginning with instruction fine-tuning, followed by a cold-start phase to elicit reasoning capabilities. Critically, we then implement reinforcement learning with grounded rewards to meticulously refine the model, aligning both its diagnostic outputs and its reasoning pathways with clinical plausibility. Quantitative assessments show that DeepMedix-R1 substantially outperforms advanced FMs, achieving improvements in report generation and visual question answering tasks. We also introduce Report Arena, a novel LLM-based benchmark that ranks DeepMedix-R1 first among competing models for output quality. Most significantly, a formal review by clinical experts reveals a profound preference for DeepMedix-R1's generated reasoning over the broadly adopted Qwen2.5-VL-7B model, confirming its superior interpretability and clinical utility.
title Toward Clinically Explainable AI for Medical Diagnosis: A Foundation Model with Human-Compatible Reasoning via Reinforcement Learning
topic Artificial Intelligence
url https://arxiv.org/abs/2509.03906