Med-R2: An Adversarial Benchmark for Evidence-Grounded Reasoning in Medical VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Wen, Niu, Fucheng, Fan, Zhiting, Xiao, Zikai, Liu, Jiaxiang, Liu, Zuozhu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917527773249536
author Ma, Wen
Niu, Fucheng
Fan, Zhiting
Xiao, Zikai
Liu, Jiaxiang
Liu, Zuozhu
author_facet Ma, Wen
Niu, Fucheng
Fan, Zhiting
Xiao, Zikai
Liu, Jiaxiang
Liu, Zuozhu
contents Vision-language models have demonstrated impressive capabilities in general medical visual question answering, yet due to limited interpretability, it remains unclear whether their predictions reflect evidence-grounded clinical reasoning or reliance on spurious priors. We introduce Med-R2 Bench, a hierarchical benchmark aligned with the clinical workflow to evaluate adversarial robustness with visual grounding. We design stepwise QA tasks to assess whether reasoning chains are strictly grounded in visual evidence across the four clinical stages, and employ adversarial perturbations to test robustness against misleading cues. Med-R2 comprises 42,432 images, 31 task categories, and 110,406 QA pairs. Evaluation across 14 VLMs reveals a sequential performance degradation along the four-stage clinical workflow. Adversarial experiments show that models rely heavily on correct prompts to guess answers. Even when provided with explicit visual cues, the models struggle to accurately align textual descriptions. Finally, we demonstrate stepwise fine-tuning using our hierarchical data significantly improves reasoning robustness, highlighting its potential to drive future improvements in evidence-based medical AI.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24492
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Med-R2: An Adversarial Benchmark for Evidence-Grounded Reasoning in Medical VLMs
Ma, Wen
Niu, Fucheng
Fan, Zhiting
Xiao, Zikai
Liu, Jiaxiang
Liu, Zuozhu
Computer Vision and Pattern Recognition
Vision-language models have demonstrated impressive capabilities in general medical visual question answering, yet due to limited interpretability, it remains unclear whether their predictions reflect evidence-grounded clinical reasoning or reliance on spurious priors. We introduce Med-R2 Bench, a hierarchical benchmark aligned with the clinical workflow to evaluate adversarial robustness with visual grounding. We design stepwise QA tasks to assess whether reasoning chains are strictly grounded in visual evidence across the four clinical stages, and employ adversarial perturbations to test robustness against misleading cues. Med-R2 comprises 42,432 images, 31 task categories, and 110,406 QA pairs. Evaluation across 14 VLMs reveals a sequential performance degradation along the four-stage clinical workflow. Adversarial experiments show that models rely heavily on correct prompts to guess answers. Even when provided with explicit visual cues, the models struggle to accurately align textual descriptions. Finally, we demonstrate stepwise fine-tuning using our hierarchical data significantly improves reasoning robustness, highlighting its potential to drive future improvements in evidence-based medical AI.
title Med-R2: An Adversarial Benchmark for Evidence-Grounded Reasoning in Medical VLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.24492