Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Murlidaran, Shravan, Wen, Ziqi, Shehabi, Sana, Eckstein, Miguel P. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps
by: Wen, Ziqi, et al.
Published: (2025)
by: Wen, Ziqi, et al.
Published: (2025)
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
by: Wen, Ziqi, et al.
Published: (2026)
by: Wen, Ziqi, et al.
Published: (2026)
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
by: Fuller, Anthony, et al.
Published: (2025)
by: Fuller, Anthony, et al.
Published: (2025)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
by: Shen, Yuxiang, et al.
Published: (2026)
by: Shen, Yuxiang, et al.
Published: (2026)
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
by: Chuang, Ian, et al.
Published: (2025)
by: Chuang, Ian, et al.
Published: (2025)
Concept-Based Explanations in Computer Vision: Where Are We and Where Could We Go?
by: Lee, Jae Hee, et al.
Published: (2024)
by: Lee, Jae Hee, et al.
Published: (2024)
Pay Attention to Where You Looked
by: Berian, Alex, et al.
Published: (2026)
by: Berian, Alex, et al.
Published: (2026)
Where do Large Vision-Language Models Look at when Answering Questions?
by: Xing, Xiaoying, et al.
Published: (2025)
by: Xing, Xiaoying, et al.
Published: (2025)
Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild
by: Teney, Damien, et al.
Published: (2025)
by: Teney, Damien, et al.
Published: (2025)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
by: Duan, Yuxiang, et al.
Published: (2025)
by: Duan, Yuxiang, et al.
Published: (2025)
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
by: Min, Cheolhong, et al.
Published: (2026)
by: Min, Cheolhong, et al.
Published: (2026)
Can We Go Beyond Visual Features? Neural Tissue Relation Modeling for Relational Graph Analysis in Non-Melanoma Skin Histology
by: Venkatraman, Shravan, et al.
Published: (2025)
by: Venkatraman, Shravan, et al.
Published: (2025)
DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
by: Park, Sungjune, et al.
Published: (2025)
by: Park, Sungjune, et al.
Published: (2025)
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
by: Lu, Jinda, et al.
Published: (2026)
by: Lu, Jinda, et al.
Published: (2026)
DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
by: Skaza, Jonathan, et al.
Published: (2025)
by: Skaza, Jonathan, et al.
Published: (2025)
ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe
by: Bai, Yifan, et al.
Published: (2023)
by: Bai, Yifan, et al.
Published: (2023)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models
by: Min, Juhong, et al.
Published: (2026)
by: Min, Juhong, et al.
Published: (2026)
Why Does It Look There? Structured Explanations for Image Classification
by: Li, Jiarui, et al.
Published: (2026)
by: Li, Jiarui, et al.
Published: (2026)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2025)
by: Zhang, Jiarui, et al.
Published: (2025)
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
by: Shabtay, Nimrod, et al.
Published: (2026)
by: Shabtay, Nimrod, et al.
Published: (2026)
INTERLACE: Interleaved Layer Pruning and Efficient Adaptation in Large Vision-Language Models
by: Madinei, Parsa, et al.
Published: (2025)
by: Madinei, Parsa, et al.
Published: (2025)
LookCloser: Frequency-aware Radiance Field for Tiny-Detail Scene
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
Where We Have Arrived in Proving the Emergence of Sparse Symbolic Concepts in AI Models
by: Ren, Qihan, et al.
Published: (2023)
by: Ren, Qihan, et al.
Published: (2023)
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
by: Salamatian, Ali, et al.
Published: (2026)
by: Salamatian, Ali, et al.
Published: (2026)
Towards Understanding Depth Perception in Foveated Rendering
by: Kergaßner, Sophie, et al.
Published: (2025)
by: Kergaßner, Sophie, et al.
Published: (2025)
Learning to Look before Learning to Like: Incorporating Human Visual Cognition into Aesthetic Quality Assessment
by: Yu, Liwen, et al.
Published: (2026)
by: Yu, Liwen, et al.
Published: (2026)
Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
by: Yang, Jiashu, et al.
Published: (2025)
by: Yang, Jiashu, et al.
Published: (2025)
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT
by: Ging, Simon, et al.
Published: (2026)
by: Ging, Simon, et al.
Published: (2026)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
Visual Acuity Consistent Foveated Rendering towards Retinal Resolution
by: Zhang, Zhi, et al.
Published: (2025)
by: Zhang, Zhi, et al.
Published: (2025)
Where Do We Stand with Implicit Neural Representations? A Technical and Performance Survey
by: Essakine, Amer, et al.
Published: (2024)
by: Essakine, Amer, et al.
Published: (2024)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
by: Lai, Bolin, et al.
Published: (2023)
by: Lai, Bolin, et al.
Published: (2023)
Learning Where to Look: Self-supervised Viewpoint Selection for Active Localization using Geometrical Information
by: Di Giammarino, Luca, et al.
Published: (2024)
by: Di Giammarino, Luca, et al.
Published: (2024)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
by: Jian, Pu, et al.
Published: (2025)
by: Jian, Pu, et al.
Published: (2025)
Have We Scene It All? Scene Graph-Aware Deep Point Cloud Compression
by: Stathoulopoulos, Nikolaos, et al.
Published: (2025)
by: Stathoulopoulos, Nikolaos, et al.
Published: (2025)
Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow Matching
by: Hu, Xin, et al.
Published: (2026)
by: Hu, Xin, et al.
Published: (2026)
Look Hear: Gaze Prediction for Speech-directed Human Attention
by: Mondal, Sounak, et al.
Published: (2024)
by: Mondal, Sounak, et al.
Published: (2024)
Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation
by: Kim, Yunsoo, et al.
Published: (2025)
by: Kim, Yunsoo, et al.
Published: (2025)
Similar Items
-
Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps
by: Wen, Ziqi, et al.
Published: (2025) -
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
by: Wen, Ziqi, et al.
Published: (2026) -
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
by: Fuller, Anthony, et al.
Published: (2025) -
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
by: Shen, Yuxiang, et al.
Published: (2026) -
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
by: Chuang, Ian, et al.
Published: (2025)