Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khan, Sumra, Chhabriya, Sagar, Zafar, Aizan, Arif, Sheeraz, Muneer, Amgad, Zafar, Anas, Raza, Shaina, Qureshi, Rizwan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910118345441280
author Khan, Sumra
Chhabriya, Sagar
Zafar, Aizan
Arif, Sheeraz
Muneer, Amgad
Zafar, Anas
Raza, Shaina
Qureshi, Rizwan
author_facet Khan, Sumra
Chhabriya, Sagar
Zafar, Aizan
Arif, Sheeraz
Muneer, Amgad
Zafar, Anas
Raza, Shaina
Qureshi, Rizwan
contents Medical vision-language models (VLMs) show strong performance on radiology tasks but often produce fluent yet weakly grounded conclusions due to over-reliance on a dominant modality. We introduce a context-aligned reasoning framework that enforces agreement across heterogeneous clinical evidence before generating diagnostic conclusions. The proposed approach augments a frozen VLM with structured contextual signals derived from radiomic statistics, explainability activations, and vocabulary-grounded semantic cues. Instead of producing free-form responses, the model generates structured outputs containing supporting evidence, uncertainty estimates, limitations, and safety notes. We observe that auxiliary signals alone provide limited benefit; performance gains emerge only when these signals are integrated through contextual verification. Experiments on chest X-ray datasets demonstrate that context alignment improves discriminative performance (AUC 0.918 to 0.925) while maintaining calibrated uncertainty. The framework also substantially reduces hallucinated keywords (1.14 to 0.25) and produces more concise reasoning explanations (19.4 to 15.3 words) without increasing model confidence (0.70 to 0.68). Cross-dataset evaluation on CheXpert further reveals that modality informativeness significantly influences reasoning behavior. These results suggest that enforcing multi-evidence agreement improves both reliability and trustworthiness in medical multimodal reasoning, while preserving the underlying model architecture.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08815
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models
Khan, Sumra
Chhabriya, Sagar
Zafar, Aizan
Arif, Sheeraz
Muneer, Amgad
Zafar, Anas
Raza, Shaina
Qureshi, Rizwan
Computer Vision and Pattern Recognition
Medical vision-language models (VLMs) show strong performance on radiology tasks but often produce fluent yet weakly grounded conclusions due to over-reliance on a dominant modality. We introduce a context-aligned reasoning framework that enforces agreement across heterogeneous clinical evidence before generating diagnostic conclusions. The proposed approach augments a frozen VLM with structured contextual signals derived from radiomic statistics, explainability activations, and vocabulary-grounded semantic cues. Instead of producing free-form responses, the model generates structured outputs containing supporting evidence, uncertainty estimates, limitations, and safety notes. We observe that auxiliary signals alone provide limited benefit; performance gains emerge only when these signals are integrated through contextual verification. Experiments on chest X-ray datasets demonstrate that context alignment improves discriminative performance (AUC 0.918 to 0.925) while maintaining calibrated uncertainty. The framework also substantially reduces hallucinated keywords (1.14 to 0.25) and produces more concise reasoning explanations (19.4 to 15.3 words) without increasing model confidence (0.70 to 0.68). Cross-dataset evaluation on CheXpert further reveals that modality informativeness significantly influences reasoning behavior. These results suggest that enforcing multi-evidence agreement improves both reliability and trustworthiness in medical multimodal reasoning, while preserving the underlying model architecture.
title Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.08815