Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910044055928832 |
|---|---|
| author | Li, Yiwei Wu, Zihao Lv, Yanjun Jiang, Hanqi You, Weihang Liu, Zhengliang Zhu, Dajiang Li, Xiang Li, Quanzheng Liu, Tianming Zhao, Lin |
| author_facet | Li, Yiwei Wu, Zihao Lv, Yanjun Jiang, Hanqi You, Weihang Liu, Zhengliang Zhu, Dajiang Li, Xiang Li, Quanzheng Liu, Tianming Zhao, Lin |
| contents | Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time. We use eye-gaze as supervision to guide VLM reasoning by introducing a small set of dedicated gaze tokens. These tokens are trained to predict gaze-selected image patch indices in temporal order, encouraging the model to follow human-like evidence acquisition and integration. Experiments on MIMIC-EYE and multiple external zero-shot benchmarks show consistent gains over baselines, achieving state-of-the-art in-domain performance and improved out-of-domain robustness. These results highlight temporally ordered gaze as an effective supervision signal for learning visually grounded medical reasoning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_06697 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs Li, Yiwei Wu, Zihao Lv, Yanjun Jiang, Hanqi You, Weihang Liu, Zhengliang Zhu, Dajiang Li, Xiang Li, Quanzheng Liu, Tianming Zhao, Lin Computer Vision and Pattern Recognition Artificial Intelligence Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time. We use eye-gaze as supervision to guide VLM reasoning by introducing a small set of dedicated gaze tokens. These tokens are trained to predict gaze-selected image patch indices in temporal order, encouraging the model to follow human-like evidence acquisition and integration. Experiments on MIMIC-EYE and multiple external zero-shot benchmarks show consistent gains over baselines, achieving state-of-the-art in-domain performance and improved out-of-domain robustness. These results highlight temporally ordered gaze as an effective supervision signal for learning visually grounded medical reasoning. |
| title | Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2603.06697 |