In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Peng, Taiying, Hua, Jiacheng, Liu, Miao, Lu, Feng |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
par: Lai, Bolin, et autres
Publié: (2022)
par: Lai, Bolin, et autres
Publié: (2022)
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
par: Zhang, Haoyu, et autres
Publié: (2025)
par: Zhang, Haoyu, et autres
Publié: (2025)
EyeCue: Driver Cognitive Distraction Detection via Gaze-Empowered Egocentric Video Understanding
par: Zhang, Lang, et autres
Publié: (2026)
par: Zhang, Lang, et autres
Publié: (2026)
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
par: Lall, Vishakha, et autres
Publié: (2025)
par: Lall, Vishakha, et autres
Publié: (2025)
EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next
par: Pan, Ye, et autres
Publié: (2026)
par: Pan, Ye, et autres
Publié: (2026)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
par: Mazzamuto, Michele, et autres
Publié: (2024)
par: Mazzamuto, Michele, et autres
Publié: (2024)
EgoCampus: Egocentric Pedestrian Eye Gaze Model and Dataset
par: John, Ronan, et autres
Publié: (2025)
par: John, Ronan, et autres
Publié: (2025)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
par: Zhou, Wenqi, et autres
Publié: (2025)
par: Zhou, Wenqi, et autres
Publié: (2025)
Gaze-Guided 3D Hand Motion Prediction for Detecting Intent in Egocentric Grasping Tasks
par: He, Yufei, et autres
Publié: (2025)
par: He, Yufei, et autres
Publié: (2025)
Personalized Federated Learning for Egocentric Video Gaze Estimation with Comprehensive Parameter Frezzing
par: Feng, Yuhu, et autres
Publié: (2025)
par: Feng, Yuhu, et autres
Publié: (2025)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
par: Lai, Bolin, et autres
Publié: (2023)
par: Lai, Bolin, et autres
Publié: (2023)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
par: Zhu, Bingwen, et autres
Publié: (2026)
par: Zhu, Bingwen, et autres
Publié: (2026)
Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
par: Liu, Shaonan, et autres
Publié: (2026)
par: Liu, Shaonan, et autres
Publié: (2026)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
par: Pani, Anupam, et autres
Publié: (2025)
par: Pani, Anupam, et autres
Publié: (2025)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
par: Seth, Ashish, et autres
Publié: (2025)
par: Seth, Ashish, et autres
Publié: (2025)
Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models
par: Wang, Hengfei, et autres
Publié: (2026)
par: Wang, Hengfei, et autres
Publié: (2026)
Real Eyes Realize Faster: Gaze Stability and Pupil Novelty for Efficient Egocentric Learning
par: Subramanian, Ajan, et autres
Publié: (2026)
par: Subramanian, Ajan, et autres
Publié: (2026)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
par: Xu, Zhiyang, et autres
Publié: (2026)
par: Xu, Zhiyang, et autres
Publié: (2026)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
par: Lee, Daeun, et autres
Publié: (2025)
par: Lee, Daeun, et autres
Publié: (2025)
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
par: Xun, Shuhang, et autres
Publié: (2025)
par: Xun, Shuhang, et autres
Publié: (2025)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
par: Chen, Boyu, et autres
Publié: (2025)
par: Chen, Boyu, et autres
Publié: (2025)
LG-Gaze: Learning Geometry-aware Continuous Prompts for Language-Guided Gaze Estimation
par: Yin, Pengwei, et autres
Publié: (2024)
par: Yin, Pengwei, et autres
Publié: (2024)
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
par: Salamatian, Ali, et autres
Publié: (2025)
par: Salamatian, Ali, et autres
Publié: (2025)
Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos
par: Materia, Daniele, et autres
Publié: (2026)
par: Materia, Daniele, et autres
Publié: (2026)
Listening with the Eyes: Benchmarking Egocentric Co-Speech Grounding across Space and Time
par: Zhou, Weijie, et autres
Publié: (2026)
par: Zhou, Weijie, et autres
Publié: (2026)
Egocentric Gaze Estimation via Neck-Mounted Camera
par: Huang, Haoyu, et autres
Publié: (2026)
par: Huang, Haoyu, et autres
Publié: (2026)
ARGaze: Autoregressive Transformers for Online Egocentric Gaze Estimation
par: Li, Jia, et autres
Publié: (2026)
par: Li, Jia, et autres
Publié: (2026)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
par: Plizzari, Chiara, et autres
Publié: (2025)
par: Plizzari, Chiara, et autres
Publié: (2025)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
par: Yin, Yufei, et autres
Publié: (2026)
par: Yin, Yufei, et autres
Publié: (2026)
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
par: Kumar, Prajneya, et autres
Publié: (2023)
par: Kumar, Prajneya, et autres
Publié: (2023)
Efficient Motion-Aware Video MLLM
par: Zhao, Zijia, et autres
Publié: (2025)
par: Zhao, Zijia, et autres
Publié: (2025)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
par: Wang, Kaibin, et autres
Publié: (2025)
par: Wang, Kaibin, et autres
Publié: (2025)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
par: Fu, Honghao, et autres
Publié: (2026)
par: Fu, Honghao, et autres
Publié: (2026)
HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts
par: Liu, Xinyu, et autres
Publié: (2024)
par: Liu, Xinyu, et autres
Publié: (2024)
Eye Gaze Tells You Where to Compute: Gaze-Driven Efficient VLMs
par: Chen, Qinyu, et autres
Publié: (2025)
par: Chen, Qinyu, et autres
Publié: (2025)
Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
par: Liu, Zuyan, et autres
Publié: (2024)
par: Liu, Zuyan, et autres
Publié: (2024)
Appearance-based Gaze Estimation With Deep Learning: A Review and Benchmark
par: Cheng, Yihua, et autres
Publié: (2021)
par: Cheng, Yihua, et autres
Publié: (2021)
GazeCLIP: Gaze-Guided CLIP with Adaptive-Enhanced Fine-Grained Language Prompt for Deepfake Attribution and Detection
par: Zhang, Yaning, et autres
Publié: (2026)
par: Zhang, Yaning, et autres
Publié: (2026)
GazeGaussian: High-Fidelity Gaze Redirection with 3D Gaussian Splatting
par: Wei, Xiaobao, et autres
Publié: (2024)
par: Wei, Xiaobao, et autres
Publié: (2024)
EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning
par: Wang, Zeyu, et autres
Publié: (2026)
par: Wang, Zeyu, et autres
Publié: (2026)
Documents similaires
-
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
par: Lai, Bolin, et autres
Publié: (2022) -
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
par: Zhang, Haoyu, et autres
Publié: (2025) -
EyeCue: Driver Cognitive Distraction Detection via Gaze-Empowered Egocentric Video Understanding
par: Zhang, Lang, et autres
Publié: (2026) -
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
par: Lall, Vishakha, et autres
Publié: (2025) -
EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next
par: Pan, Ye, et autres
Publié: (2026)