Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Murlidaran, Shravan, Wen, Ziqi, Shehabi, Sana, Eckstein, Miguel P. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps
di: Wen, Ziqi, et al.
Pubblicazione: (2025)
di: Wen, Ziqi, et al.
Pubblicazione: (2025)
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
di: Wen, Ziqi, et al.
Pubblicazione: (2026)
di: Wen, Ziqi, et al.
Pubblicazione: (2026)
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
di: Fuller, Anthony, et al.
Pubblicazione: (2025)
di: Fuller, Anthony, et al.
Pubblicazione: (2025)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
di: Shen, Yuxiang, et al.
Pubblicazione: (2026)
di: Shen, Yuxiang, et al.
Pubblicazione: (2026)
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
di: Chuang, Ian, et al.
Pubblicazione: (2025)
di: Chuang, Ian, et al.
Pubblicazione: (2025)
Concept-Based Explanations in Computer Vision: Where Are We and Where Could We Go?
di: Lee, Jae Hee, et al.
Pubblicazione: (2024)
di: Lee, Jae Hee, et al.
Pubblicazione: (2024)
Pay Attention to Where You Looked
di: Berian, Alex, et al.
Pubblicazione: (2026)
di: Berian, Alex, et al.
Pubblicazione: (2026)
Where do Large Vision-Language Models Look at when Answering Questions?
di: Xing, Xiaoying, et al.
Pubblicazione: (2025)
di: Xing, Xiaoying, et al.
Pubblicazione: (2025)
Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild
di: Teney, Damien, et al.
Pubblicazione: (2025)
di: Teney, Damien, et al.
Pubblicazione: (2025)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
di: Duan, Yuxiang, et al.
Pubblicazione: (2025)
di: Duan, Yuxiang, et al.
Pubblicazione: (2025)
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
di: Min, Cheolhong, et al.
Pubblicazione: (2026)
di: Min, Cheolhong, et al.
Pubblicazione: (2026)
Can We Go Beyond Visual Features? Neural Tissue Relation Modeling for Relational Graph Analysis in Non-Melanoma Skin Histology
di: Venkatraman, Shravan, et al.
Pubblicazione: (2025)
di: Venkatraman, Shravan, et al.
Pubblicazione: (2025)
DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
di: Park, Sungjune, et al.
Pubblicazione: (2025)
di: Park, Sungjune, et al.
Pubblicazione: (2025)
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
di: Lu, Jinda, et al.
Pubblicazione: (2026)
di: Lu, Jinda, et al.
Pubblicazione: (2026)
DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
di: Skaza, Jonathan, et al.
Pubblicazione: (2025)
di: Skaza, Jonathan, et al.
Pubblicazione: (2025)
ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe
di: Bai, Yifan, et al.
Pubblicazione: (2023)
di: Bai, Yifan, et al.
Pubblicazione: (2023)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models
di: Min, Juhong, et al.
Pubblicazione: (2026)
di: Min, Juhong, et al.
Pubblicazione: (2026)
Why Does It Look There? Structured Explanations for Image Classification
di: Li, Jiarui, et al.
Pubblicazione: (2026)
di: Li, Jiarui, et al.
Pubblicazione: (2026)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
di: Zhang, Jiarui, et al.
Pubblicazione: (2025)
di: Zhang, Jiarui, et al.
Pubblicazione: (2025)
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
di: Shabtay, Nimrod, et al.
Pubblicazione: (2026)
di: Shabtay, Nimrod, et al.
Pubblicazione: (2026)
INTERLACE: Interleaved Layer Pruning and Efficient Adaptation in Large Vision-Language Models
di: Madinei, Parsa, et al.
Pubblicazione: (2025)
di: Madinei, Parsa, et al.
Pubblicazione: (2025)
LookCloser: Frequency-aware Radiance Field for Tiny-Detail Scene
di: Zhang, Xiaoyu, et al.
Pubblicazione: (2025)
di: Zhang, Xiaoyu, et al.
Pubblicazione: (2025)
Where We Have Arrived in Proving the Emergence of Sparse Symbolic Concepts in AI Models
di: Ren, Qihan, et al.
Pubblicazione: (2023)
di: Ren, Qihan, et al.
Pubblicazione: (2023)
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
di: Salamatian, Ali, et al.
Pubblicazione: (2026)
di: Salamatian, Ali, et al.
Pubblicazione: (2026)
Towards Understanding Depth Perception in Foveated Rendering
di: Kergaßner, Sophie, et al.
Pubblicazione: (2025)
di: Kergaßner, Sophie, et al.
Pubblicazione: (2025)
Learning to Look before Learning to Like: Incorporating Human Visual Cognition into Aesthetic Quality Assessment
di: Yu, Liwen, et al.
Pubblicazione: (2026)
di: Yu, Liwen, et al.
Pubblicazione: (2026)
Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
di: Yang, Jiashu, et al.
Pubblicazione: (2025)
di: Yang, Jiashu, et al.
Pubblicazione: (2025)
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
di: Li, Liyang, et al.
Pubblicazione: (2026)
di: Li, Liyang, et al.
Pubblicazione: (2026)
Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT
di: Ging, Simon, et al.
Pubblicazione: (2026)
di: Ging, Simon, et al.
Pubblicazione: (2026)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
di: Zhu, Xingyu, et al.
Pubblicazione: (2026)
di: Zhu, Xingyu, et al.
Pubblicazione: (2026)
Visual Acuity Consistent Foveated Rendering towards Retinal Resolution
di: Zhang, Zhi, et al.
Pubblicazione: (2025)
di: Zhang, Zhi, et al.
Pubblicazione: (2025)
Where Do We Stand with Implicit Neural Representations? A Technical and Performance Survey
di: Essakine, Amer, et al.
Pubblicazione: (2024)
di: Essakine, Amer, et al.
Pubblicazione: (2024)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
di: Lai, Bolin, et al.
Pubblicazione: (2023)
di: Lai, Bolin, et al.
Pubblicazione: (2023)
Learning Where to Look: Self-supervised Viewpoint Selection for Active Localization using Geometrical Information
di: Di Giammarino, Luca, et al.
Pubblicazione: (2024)
di: Di Giammarino, Luca, et al.
Pubblicazione: (2024)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
di: Jian, Pu, et al.
Pubblicazione: (2025)
di: Jian, Pu, et al.
Pubblicazione: (2025)
Have We Scene It All? Scene Graph-Aware Deep Point Cloud Compression
di: Stathoulopoulos, Nikolaos, et al.
Pubblicazione: (2025)
di: Stathoulopoulos, Nikolaos, et al.
Pubblicazione: (2025)
Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow Matching
di: Hu, Xin, et al.
Pubblicazione: (2026)
di: Hu, Xin, et al.
Pubblicazione: (2026)
Look Hear: Gaze Prediction for Speech-directed Human Attention
di: Mondal, Sounak, et al.
Pubblicazione: (2024)
di: Mondal, Sounak, et al.
Pubblicazione: (2024)
Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation
di: Kim, Yunsoo, et al.
Pubblicazione: (2025)
di: Kim, Yunsoo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps
di: Wen, Ziqi, et al.
Pubblicazione: (2025) -
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
di: Wen, Ziqi, et al.
Pubblicazione: (2026) -
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
di: Fuller, Anthony, et al.
Pubblicazione: (2025) -
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
di: Shen, Yuxiang, et al.
Pubblicazione: (2026) -
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
di: Chuang, Ian, et al.
Pubblicazione: (2025)