GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mathew, Athul M., Hermassi, Haithem, Khalid, Thariq, Khan, Arshad Ali |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Multi-Modal Saliency and Fusion for Gaze Target Detection
von: Mathew, Athul M., et al.
Veröffentlicht: (2025)
von: Mathew, Athul M., et al.
Veröffentlicht: (2025)
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
von: Ebouky, Brown, et al.
Veröffentlicht: (2026)
von: Ebouky, Brown, et al.
Veröffentlicht: (2026)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
von: Pani, Anupam, et al.
Veröffentlicht: (2025)
von: Pani, Anupam, et al.
Veröffentlicht: (2025)
GazeQwen: Lightweight Gaze-Conditioned LLM Modulation for Streaming Video Understanding
von: Pham, Trong Thang, et al.
Veröffentlicht: (2026)
von: Pham, Trong Thang, et al.
Veröffentlicht: (2026)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
GazeMoE: Perception of Gaze Target with Mixture-of-Experts
von: Dai, Zhuangzhuang, et al.
Veröffentlicht: (2026)
von: Dai, Zhuangzhuang, et al.
Veröffentlicht: (2026)
Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models
von: Wang, Hengfei, et al.
Veröffentlicht: (2026)
von: Wang, Hengfei, et al.
Veröffentlicht: (2026)
TPP-Gaze: Modelling Gaze Dynamics in Space and Time with Neural Temporal Point Processes
von: D'Amelio, Alessandro, et al.
Veröffentlicht: (2024)
von: D'Amelio, Alessandro, et al.
Veröffentlicht: (2024)
Voila-A: Aligning Vision-Language Models with User's Gaze Attention
von: Yan, Kun, et al.
Veröffentlicht: (2023)
von: Yan, Kun, et al.
Veröffentlicht: (2023)
Seeing Like Radiologists: Context- and Gaze-Guided Vision-Language Pretraining for Chest X-rays
von: Liu, Kang, et al.
Veröffentlicht: (2026)
von: Liu, Kang, et al.
Veröffentlicht: (2026)
GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer
von: Zhao, Xinyuan, et al.
Veröffentlicht: (2026)
von: Zhao, Xinyuan, et al.
Veröffentlicht: (2026)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
Learning Gaze-aware Compositional GAN
von: Aranjuelo, Nerea, et al.
Veröffentlicht: (2024)
von: Aranjuelo, Nerea, et al.
Veröffentlicht: (2024)
DMAGaze: Gaze Estimation Based on Feature Disentanglement and Multi-Scale Attention
von: Chen, Haohan, et al.
Veröffentlicht: (2025)
von: Chen, Haohan, et al.
Veröffentlicht: (2025)
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
von: Chen, Pingyi, et al.
Veröffentlicht: (2025)
von: Chen, Pingyi, et al.
Veröffentlicht: (2025)
GazeSearch: Radiology Findings Search Benchmark
von: Pham, Trong Thang, et al.
Veröffentlicht: (2024)
von: Pham, Trong Thang, et al.
Veröffentlicht: (2024)
Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
von: Cartella, Giuseppe, et al.
Veröffentlicht: (2025)
von: Cartella, Giuseppe, et al.
Veröffentlicht: (2025)
Enhancing Human-Computer Interaction in Chest X-ray Analysis using Vision and Language Model with Eye Gaze Patterns
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
RGBD Gaze Tracking Using Transformer for Feature Fusion
von: Bauer, Tobias J.
Veröffentlicht: (2025)
von: Bauer, Tobias J.
Veröffentlicht: (2025)
Weakly-supervised Medical Image Segmentation with Gaze Annotations
von: Zhong, Yuan, et al.
Veröffentlicht: (2024)
von: Zhong, Yuan, et al.
Veröffentlicht: (2024)
3D Gaussian and Diffusion-Based Gaze Redirection
von: Panchalingam, Abiram, et al.
Veröffentlicht: (2025)
von: Panchalingam, Abiram, et al.
Veröffentlicht: (2025)
Is Geometry Enough? An Evaluation of Landmark-Based Gaze Estimation
von: Agostinelli, Daniele, et al.
Veröffentlicht: (2026)
von: Agostinelli, Daniele, et al.
Veröffentlicht: (2026)
Toward Gaze Target Detection of Young Autistic Children
von: Deng, Shijian, et al.
Veröffentlicht: (2025)
von: Deng, Shijian, et al.
Veröffentlicht: (2025)
Modeling Subjective Urban Perception with Human Gaze
von: Che, Lin, et al.
Veröffentlicht: (2026)
von: Che, Lin, et al.
Veröffentlicht: (2026)
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
von: Wang, Shijing, et al.
Veröffentlicht: (2025)
von: Wang, Shijing, et al.
Veröffentlicht: (2025)
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
Gaze-Informed Vision Transformers: Predicting Driving Decisions Under Uncertainty
von: Koorathota, Sharath, et al.
Veröffentlicht: (2023)
von: Koorathota, Sharath, et al.
Veröffentlicht: (2023)
BehaviorVLM: Unified Finetuning-Free Behavioral Understanding with Vision-Language Reasoning
von: Ke, Jingyang, et al.
Veröffentlicht: (2026)
von: Ke, Jingyang, et al.
Veröffentlicht: (2026)
GazeFusion: Saliency-Guided Image Generation
von: Zhang, Yunxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Yunxiang, et al.
Veröffentlicht: (2024)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
See Through the Noise: Improving Domain Generalization in Gaze Estimation
von: Peng, Yanming, et al.
Veröffentlicht: (2026)
von: Peng, Yanming, et al.
Veröffentlicht: (2026)
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
von: Lall, Vishakha, et al.
Veröffentlicht: (2025)
von: Lall, Vishakha, et al.
Veröffentlicht: (2025)
Unveiling the Truth: Exploring Human Gaze Patterns in Fake Images
von: Cartella, Giuseppe, et al.
Veröffentlicht: (2024)
von: Cartella, Giuseppe, et al.
Veröffentlicht: (2024)
Gaze-Guided 3D Hand Motion Prediction for Detecting Intent in Egocentric Grasping Tasks
von: He, Yufei, et al.
Veröffentlicht: (2025)
von: He, Yufei, et al.
Veröffentlicht: (2025)
PoseGaze-AHP: A Knowledge-Based 3D Dataset for AI-Driven Ocular and Postural Diagnosis
von: Al-Dabet, Saja, et al.
Veröffentlicht: (2025)
von: Al-Dabet, Saja, et al.
Veröffentlicht: (2025)
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
Toddlers' Active Gaze Behavior Supports Self-Supervised Object Learning
von: Yu, Zhengyang, et al.
Veröffentlicht: (2024)
von: Yu, Zhengyang, et al.
Veröffentlicht: (2024)
A Generalized Label Shift Perspective for Cross-Domain Gaze Estimation
von: Yang, Hao-Ran, et al.
Veröffentlicht: (2025)
von: Yang, Hao-Ran, et al.
Veröffentlicht: (2025)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
HPE-CogVLM: Advancing Vision Language Models with a Head Pose Grounding Task
von: Tian, Yu, et al.
Veröffentlicht: (2024)
von: Tian, Yu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Leveraging Multi-Modal Saliency and Fusion for Gaze Target Detection
von: Mathew, Athul M., et al.
Veröffentlicht: (2025) -
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
von: Ebouky, Brown, et al.
Veröffentlicht: (2026) -
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
von: Pani, Anupam, et al.
Veröffentlicht: (2025) -
GazeQwen: Lightweight Gaze-Conditioned LLM Modulation for Streaming Video Understanding
von: Pham, Trong Thang, et al.
Veröffentlicht: (2026) -
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
von: Lee, Daeun, et al.
Veröffentlicht: (2025)