Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yiwei, Wu, Zihao, Lv, Yanjun, Jiang, Hanqi, You, Weihang, Liu, Zhengliang, Zhu, Dajiang, Li, Xiang, Li, Quanzheng, Liu, Tianming, Zhao, Lin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910044055928832
author Li, Yiwei
Wu, Zihao
Lv, Yanjun
Jiang, Hanqi
You, Weihang
Liu, Zhengliang
Zhu, Dajiang
Li, Xiang
Li, Quanzheng
Liu, Tianming
Zhao, Lin
author_facet Li, Yiwei
Wu, Zihao
Lv, Yanjun
Jiang, Hanqi
You, Weihang
Liu, Zhengliang
Zhu, Dajiang
Li, Xiang
Li, Quanzheng
Liu, Tianming
Zhao, Lin
contents Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time. We use eye-gaze as supervision to guide VLM reasoning by introducing a small set of dedicated gaze tokens. These tokens are trained to predict gaze-selected image patch indices in temporal order, encouraging the model to follow human-like evidence acquisition and integration. Experiments on MIMIC-EYE and multiple external zero-shot benchmarks show consistent gains over baselines, achieving state-of-the-art in-domain performance and improved out-of-domain robustness. These results highlight temporally ordered gaze as an effective supervision signal for learning visually grounded medical reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06697
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
Li, Yiwei
Wu, Zihao
Lv, Yanjun
Jiang, Hanqi
You, Weihang
Liu, Zhengliang
Zhu, Dajiang
Li, Xiang
Li, Quanzheng
Liu, Tianming
Zhao, Lin
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time. We use eye-gaze as supervision to guide VLM reasoning by introducing a small set of dedicated gaze tokens. These tokens are trained to predict gaze-selected image patch indices in temporal order, encouraging the model to follow human-like evidence acquisition and integration. Experiments on MIMIC-EYE and multiple external zero-shot benchmarks show consistent gains over baselines, achieving state-of-the-art in-domain performance and improved out-of-domain robustness. These results highlight temporally ordered gaze as an effective supervision signal for learning visually grounded medical reasoning.
title Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.06697