Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Prajneya, Khandelwal, Eshika, Tapaswi, Makarand, Sreekumar, Vishnu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025)
by: Kala, Divy, et al.
Published: (2025)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
by: Khandelwal, Eshika, et al.
Published: (2025)
by: Khandelwal, Eshika, et al.
Published: (2025)
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
by: Darur, Balaji, et al.
Published: (2026)
by: Darur, Balaji, et al.
Published: (2026)
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
by: Singh, Darshan, et al.
Published: (2024)
by: Singh, Darshan, et al.
Published: (2024)
"Previously on ..." From Recaps to Story Summarization
by: Singh, Aditya Kumar, et al.
Published: (2024)
by: Singh, Aditya Kumar, et al.
Published: (2024)
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
by: Gaur, Manu, et al.
Published: (2024)
by: Gaur, Manu, et al.
Published: (2024)
MALeR: Improving Compositional Fidelity in Layout-Guided Generation
by: Saxena, Shivank, et al.
Published: (2025)
by: Saxena, Shivank, et al.
Published: (2025)
Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation
by: Gaur, Manu, et al.
Published: (2024)
by: Gaur, Manu, et al.
Published: (2024)
Investigating Mechanisms for In-Context Vision Language Binding
by: Saravanan, Darshana, et al.
Published: (2025)
by: Saravanan, Darshana, et al.
Published: (2025)
MICap: A Unified Model for Identity-aware Movie Descriptions
by: Raajesh, Haran, et al.
Published: (2024)
by: Raajesh, Haran, et al.
Published: (2024)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
by: Mazzamuto, Michele, et al.
Published: (2024)
by: Mazzamuto, Michele, et al.
Published: (2024)
Seeing Eye to AI: Human Alignment via Gaze-Based Response Rewards for Large Language Models
by: Lopez-Cardona, Angela, et al.
Published: (2024)
by: Lopez-Cardona, Angela, et al.
Published: (2024)
STRinGS: Selective Text Refinement in Gaussian Splatting
by: Raundhal, Abhinav, et al.
Published: (2025)
by: Raundhal, Abhinav, et al.
Published: (2025)
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
by: Saravanan, Darshana, et al.
Published: (2024)
by: Saravanan, Darshana, et al.
Published: (2024)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026)
by: Song, Junha, et al.
Published: (2026)
Eye Gaze as a Signal for Conveying User Attention in Contextual AI Systems
by: Wilson, Ethan, et al.
Published: (2025)
by: Wilson, Ethan, et al.
Published: (2025)
DHECA-SuperGaze: Dual Head-Eye Cross-Attention and Super-Resolution for Unconstrained Gaze Estimation
by: Šikić, Franko, et al.
Published: (2025)
by: Šikić, Franko, et al.
Published: (2025)
Steerable Visual Representations
by: Ruthardt, Jona, et al.
Published: (2026)
by: Ruthardt, Jona, et al.
Published: (2026)
SGAP-Gaze: Scene Grid Attention Based Point-of-Gaze Estimation Network for Driver Gaze
by: Sharma, Pavan Kumar, et al.
Published: (2026)
by: Sharma, Pavan Kumar, et al.
Published: (2026)
Learning to See Like Humans: Gaze-Aligned Cycling Safety Prediction
by: Perdigão, Luís Maria, et al.
Published: (2026)
by: Perdigão, Luís Maria, et al.
Published: (2026)
Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models
by: Wang, Hengfei, et al.
Published: (2026)
by: Wang, Hengfei, et al.
Published: (2026)
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
by: Peng, Taiying, et al.
Published: (2025)
by: Peng, Taiying, et al.
Published: (2025)
The Sound of Water: Inferring Physical Properties from Pouring Liquids
by: Bagad, Piyush, et al.
Published: (2024)
by: Bagad, Piyush, et al.
Published: (2024)
OASIS Uncovers: High-Quality T2I Models, Same Old Stereotypes
by: Dehdashtian, Sepehr, et al.
Published: (2025)
by: Dehdashtian, Sepehr, et al.
Published: (2025)
NurtureNet: A Multi-task Video-based Approach for Newborn Anthropometry
by: Khandelwal, Yash, et al.
Published: (2024)
by: Khandelwal, Yash, et al.
Published: (2024)
Supporting Mitosis Detection AI Training with Inter-Observer Eye-Gaze Consistencies
by: Gu, Hongyan, et al.
Published: (2024)
by: Gu, Hongyan, et al.
Published: (2024)
EgoCampus: Egocentric Pedestrian Eye Gaze Model and Dataset
by: John, Ronan, et al.
Published: (2025)
by: John, Ronan, et al.
Published: (2025)
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
by: Lall, Vishakha, et al.
Published: (2025)
by: Lall, Vishakha, et al.
Published: (2025)
Eye Gaze Tells You Where to Compute: Gaze-Driven Efficient VLMs
by: Chen, Qinyu, et al.
Published: (2025)
by: Chen, Qinyu, et al.
Published: (2025)
Foraging with the Eyes: Dynamics in Human Visual Gaze and Deep Predictive Modeling
by: Panchagnula, Tejaswi V.
Published: (2025)
by: Panchagnula, Tejaswi V.
Published: (2025)
Investigating Memorization in Video Diffusion Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Look Hear: Gaze Prediction for Speech-directed Human Attention
by: Mondal, Sounak, et al.
Published: (2024)
by: Mondal, Sounak, et al.
Published: (2024)
Eyes Tell the Truth: GazeVal Highlights Shortcomings of Generative AI in Medical Imaging
by: Wong, David, et al.
Published: (2025)
by: Wong, David, et al.
Published: (2025)
EyeCue: Driver Cognitive Distraction Detection via Gaze-Empowered Egocentric Video Understanding
by: Zhang, Lang, et al.
Published: (2026)
by: Zhang, Lang, et al.
Published: (2026)
TalkingEyes: Pluralistic Speech-Driven 3D Eye Gaze Animation
by: Zhuang, Yixiang, et al.
Published: (2025)
by: Zhuang, Yixiang, et al.
Published: (2025)
Spatio-Temporal Attention and Gaussian Processes for Personalized Video Gaze Estimation
by: Jindal, Swati, et al.
Published: (2024)
by: Jindal, Swati, et al.
Published: (2024)
Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual Foraging
by: Wang, Bo, et al.
Published: (2024)
by: Wang, Bo, et al.
Published: (2024)
Airway Skill Assessment with Spatiotemporal Attention Mechanisms Using Human Gaze
by: Ainam, Jean-Paul, et al.
Published: (2025)
by: Ainam, Jean-Paul, et al.
Published: (2025)
Seeing the World through Your Eyes
by: Alzayer, Hadi, et al.
Published: (2023)
by: Alzayer, Hadi, et al.
Published: (2023)
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
by: Shi, Baifeng, et al.
Published: (2026)
by: Shi, Baifeng, et al.
Published: (2026)
Similar Items
-
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025) -
More than a Moment: Towards Coherent Sequences of Audio Descriptions
by: Khandelwal, Eshika, et al.
Published: (2025) -
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
by: Darur, Balaji, et al.
Published: (2026) -
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
by: Singh, Darshan, et al.
Published: (2024) -
"Previously on ..." From Recaps to Story Summarization
by: Singh, Aditya Kumar, et al.
Published: (2024)