Saved in:
Bibliographic Details
Main Authors: Nguyen, Minh Khoi, Le, Dai Lam, Jafari, Amir Reza, Nguyen, Tuan Dung, Son, Mai Hong, Thong, Mai Huy, Nguyen, Quang Huy, Nguyen, Thanh Trung, Farahbakhsh, Reza, Crespi, Noel, Nguyen, Phi Le
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.10002
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909031092715520
author Nguyen, Minh Khoi
Le, Dai Lam
Jafari, Amir Reza
Nguyen, Tuan Dung
Son, Mai Hong
Thong, Mai Huy
Nguyen, Quang Huy
Nguyen, Thanh Trung
Farahbakhsh, Reza
Crespi, Noel
Nguyen, Phi Le
author_facet Nguyen, Minh Khoi
Le, Dai Lam
Jafari, Amir Reza
Nguyen, Tuan Dung
Son, Mai Hong
Thong, Mai Huy
Nguyen, Quang Huy
Nguyen, Thanh Trung
Farahbakhsh, Reza
Crespi, Noel
Nguyen, Phi Le
contents Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet incorrect statements, raising significant safety concerns. Existing medical hallucination benchmarks primarily focus on 2D imaging with one-shot diagnostic questions, offering limited insight into whether predictions are grounded in correct localization and abnormality identification, allowing critical reasoning errors to remain hidden behind seemingly correct diagnoses. We introduce Med-StepBench, the first large-scale benchmark for step-wise hallucination detection in 3D oncological PET/CT, comprising over 12,000 images and more than 1,000,000 image-statement pairs across volumetric and multi-view 2D data, which decomposes clinical reasoning into four expert-designed diagnostic stages. Using clinician-verified annotations, we perform the first step-level evaluation of general-purpose and medical VLMs, revealing systematic failure modes obscured by aggregate accuracy metrics. Furthermore, we show that current VLMs are highly susceptible to adversarial yet clinically plausible intermediate explanations, which significantly amplify hallucinations despite contradictory visual evidence. Together, our findings highlight fundamental limitations in grounding multi-step clinical reasoning and establish Med-StepBench as a rigorous benchmark for developing safer and more reliable medical VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10002
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models
Nguyen, Minh Khoi
Le, Dai Lam
Jafari, Amir Reza
Nguyen, Tuan Dung
Son, Mai Hong
Thong, Mai Huy
Nguyen, Quang Huy
Nguyen, Thanh Trung
Farahbakhsh, Reza
Crespi, Noel
Nguyen, Phi Le
Computer Vision and Pattern Recognition
Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet incorrect statements, raising significant safety concerns. Existing medical hallucination benchmarks primarily focus on 2D imaging with one-shot diagnostic questions, offering limited insight into whether predictions are grounded in correct localization and abnormality identification, allowing critical reasoning errors to remain hidden behind seemingly correct diagnoses. We introduce Med-StepBench, the first large-scale benchmark for step-wise hallucination detection in 3D oncological PET/CT, comprising over 12,000 images and more than 1,000,000 image-statement pairs across volumetric and multi-view 2D data, which decomposes clinical reasoning into four expert-designed diagnostic stages. Using clinician-verified annotations, we perform the first step-level evaluation of general-purpose and medical VLMs, revealing systematic failure modes obscured by aggregate accuracy metrics. Furthermore, we show that current VLMs are highly susceptible to adversarial yet clinically plausible intermediate explanations, which significantly amplify hallucinations despite contradictory visual evidence. Together, our findings highlight fundamental limitations in grounding multi-step clinical reasoning and establish Med-StepBench as a rigorous benchmark for developing safer and more reliable medical VLMs.
title Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.10002