Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shaonan, Yu, Guo, Luo, Xiaoling, Zheng, Shiyi, Chen, Wenting, Liu, Jie, Shen, Linlin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917194765434880
author Liu, Shaonan
Yu, Guo
Luo, Xiaoling
Zheng, Shiyi
Chen, Wenting
Liu, Jie
Shen, Linlin
author_facet Liu, Shaonan
Yu, Guo
Luo, Xiaoling
Zheng, Shiyi
Chen, Wenting
Liu, Jie
Shen, Linlin
contents Medical Multimodal Large Language Models (Med-MLLMs) require egocentric clinical intent understanding for real-world deployment, yet existing benchmarks fail to evaluate this critical capability. To address these challenges, we introduce MedGaze-Bench, the first benchmark leveraging clinician gaze as a Cognitive Cursor to assess intent understanding across surgery, emergency simulation, and diagnostic interpretation. Our benchmark addresses three fundamental challenges: visual homogeneity of anatomical structures, strict temporal-causal dependencies in clinical workflows, and implicit adherence to safety protocols. We propose a Three-Dimensional Clinical Intent Framework evaluating: (1) Spatial Intent: discriminating precise targets amid visual noise, (2) Temporal Intent: inferring causal rationale through retrospective and prospective reasoning, and (3) Standard Intent: verifying protocol compliance through safety checks. Beyond accuracy metrics, we introduce Trap QA mechanisms to stress-test clinical reliability by penalizing hallucinations and cognitive sycophancy. Experiments reveal current MLLMs struggle with egocentric intent due to over-reliance on global features, leading to fabricated observations and uncritical acceptance of invalid instructions.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06750
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
Liu, Shaonan
Yu, Guo
Luo, Xiaoling
Zheng, Shiyi
Chen, Wenting
Liu, Jie
Shen, Linlin
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Medical Multimodal Large Language Models (Med-MLLMs) require egocentric clinical intent understanding for real-world deployment, yet existing benchmarks fail to evaluate this critical capability. To address these challenges, we introduce MedGaze-Bench, the first benchmark leveraging clinician gaze as a Cognitive Cursor to assess intent understanding across surgery, emergency simulation, and diagnostic interpretation. Our benchmark addresses three fundamental challenges: visual homogeneity of anatomical structures, strict temporal-causal dependencies in clinical workflows, and implicit adherence to safety protocols. We propose a Three-Dimensional Clinical Intent Framework evaluating: (1) Spatial Intent: discriminating precise targets amid visual noise, (2) Temporal Intent: inferring causal rationale through retrospective and prospective reasoning, and (3) Standard Intent: verifying protocol compliance through safety checks. Beyond accuracy metrics, we introduce Trap QA mechanisms to stress-test clinical reliability by penalizing hallucinations and cognitive sycophancy. Experiments reveal current MLLMs struggle with egocentric intent due to over-reliance on global features, leading to fabricated observations and uncritical acceptance of invalid instructions.
title Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.06750