MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bao, Zhijie, Chen, Fangke, Bao, Licheng, Zhang, Chenhui, Chen, Wei, Peng, Jiajie, Wei, Zhongyu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910131333103616
author Bao, Zhijie
Chen, Fangke
Bao, Licheng
Zhang, Chenhui
Chen, Wei
Peng, Jiajie
Wei, Zhongyu
author_facet Bao, Zhijie
Chen, Fangke
Bao, Licheng
Zhang, Chenhui
Chen, Wei
Peng, Jiajie
Wei, Zhongyu
contents The potential of Multimodal Large Language Models (MLLMs) in domain of medical imaging raise the demands of systematic and rigorous evaluation frameworks that are aligned with the real-world medical imaging practice. Existing practices that report single or coarse-grained metrics are lack the granularity required for specialized clinical support and fail to assess the reliability of reasoning mechanisms. To address this, we propose a paradigm shift toward multidimensional, fine-grained and in-depth evaluation. Based on a two-stage systematic construction pipeline designed for this paradigm, we instantiate it with MedRCube. We benchmark 33 MLLMs, \textit{Lingshu-32B} achieve top-tier performance. Crucially, MedRCube exposes a series of pronounced insights inaccessible under prior evaluation settings. Furthermore, we introduce a credibility evaluation subset to quantify reasoning credibility, uncover a highly significant positive association between shortcut behavior and diagnostic task performance, raising concerns for clinically trustworthy deployment. The resources of this work can be found at https://github.com/F1mc/MedRCube.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13756
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging
Bao, Zhijie
Chen, Fangke
Bao, Licheng
Zhang, Chenhui
Chen, Wei
Peng, Jiajie
Wei, Zhongyu
Computation and Language
Computer Vision and Pattern Recognition
The potential of Multimodal Large Language Models (MLLMs) in domain of medical imaging raise the demands of systematic and rigorous evaluation frameworks that are aligned with the real-world medical imaging practice. Existing practices that report single or coarse-grained metrics are lack the granularity required for specialized clinical support and fail to assess the reliability of reasoning mechanisms. To address this, we propose a paradigm shift toward multidimensional, fine-grained and in-depth evaluation. Based on a two-stage systematic construction pipeline designed for this paradigm, we instantiate it with MedRCube. We benchmark 33 MLLMs, \textit{Lingshu-32B} achieve top-tier performance. Crucially, MedRCube exposes a series of pronounced insights inaccessible under prior evaluation settings. Furthermore, we introduce a credibility evaluation subset to quantify reasoning credibility, uncover a highly significant positive association between shortcut behavior and diagnostic task performance, raising concerns for clinically trustworthy deployment. The resources of this work can be found at https://github.com/F1mc/MedRCube.
title MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.13756