EMAD: Evidence-Centric Grounded Multimodal Diagnosis for Alzheimer's Disease

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Qiuhui, Yao, Xuancheng, Zhou, Zhenglei, Hu, Xinyue, Hong, Yi
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910029268910080
author Chen, Qiuhui
Yao, Xuancheng
Zhou, Zhenglei
Hu, Xinyue
Hong, Yi
author_facet Chen, Qiuhui
Yao, Xuancheng
Zhou, Zhenglei
Hu, Xinyue
Hong, Yi
contents Deep learning models for medical image analysis often act as black boxes, seldom aligning with clinical guidelines or explicitly linking decisions to supporting evidence. This is especially critical in Alzheimer's disease (AD), where predictions should be grounded in both anatomical and clinical findings. We present EMAD, a vision-language framework that generates structured AD diagnostic reports in which each claim is explicitly grounded in multimodal evidence. EMAD uses a hierarchical Sentence-Evidence-Anatomy (SEA) grounding mechanism: (i) sentence-to-evidence grounding links generated sentences to clinical evidence phrases, and (ii) evidence-to-anatomy grounding localizes corresponding structures on 3D brain MRI. To reduce dense annotation requirements, we propose GTX-Distill, which transfers grounding behavior from a teacher trained with limited supervision to a student operating on model-generated reports. We further introduce Executable-Rule GRPO, a reinforcement fine-tuning scheme with verifiable rewards that enforces clinical consistency, protocol adherence, and reasoning-diagnosis coherence. On the AD-MultiSense dataset, EMAD achieves state-of-the-art diagnostic accuracy and produces more transparent, anatomically faithful reports than existing methods. We will release code and grounding annotations to support future research in trustworthy medical vision-language models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19178
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EMAD: Evidence-Centric Grounded Multimodal Diagnosis for Alzheimer's Disease
Chen, Qiuhui
Yao, Xuancheng
Zhou, Zhenglei
Hu, Xinyue
Hong, Yi
Computer Vision and Pattern Recognition
Deep learning models for medical image analysis often act as black boxes, seldom aligning with clinical guidelines or explicitly linking decisions to supporting evidence. This is especially critical in Alzheimer's disease (AD), where predictions should be grounded in both anatomical and clinical findings. We present EMAD, a vision-language framework that generates structured AD diagnostic reports in which each claim is explicitly grounded in multimodal evidence. EMAD uses a hierarchical Sentence-Evidence-Anatomy (SEA) grounding mechanism: (i) sentence-to-evidence grounding links generated sentences to clinical evidence phrases, and (ii) evidence-to-anatomy grounding localizes corresponding structures on 3D brain MRI. To reduce dense annotation requirements, we propose GTX-Distill, which transfers grounding behavior from a teacher trained with limited supervision to a student operating on model-generated reports. We further introduce Executable-Rule GRPO, a reinforcement fine-tuning scheme with verifiable rewards that enforces clinical consistency, protocol adherence, and reasoning-diagnosis coherence. On the AD-MultiSense dataset, EMAD achieves state-of-the-art diagnostic accuracy and produces more transparent, anatomically faithful reports than existing methods. We will release code and grounding annotations to support future research in trustworthy medical vision-language models.
title EMAD: Evidence-Centric Grounded Multimodal Diagnosis for Alzheimer's Disease
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.19178