GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Ebouky, Brown, Carrino, Gabriele, Avogaro, Niccolo, Studer, Christoph, Bartezzaghi, Andrea, Rigotti, Mattia |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Eliciting Reasoning in Language Models with Cognitive Tools
by: Ebouky, Brown, et al.
Published: (2025)
by: Ebouky, Brown, et al.
Published: (2025)
Show or Tell? Effectively prompting Vision-Language Models for semantic segmentation
by: Avogaro, Niccolo, et al.
Published: (2025)
by: Avogaro, Niccolo, et al.
Published: (2025)
Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
by: Ebouky, Brown, et al.
Published: (2025)
by: Ebouky, Brown, et al.
Published: (2025)
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
by: Mathew, Athul M., et al.
Published: (2025)
by: Mathew, Athul M., et al.
Published: (2025)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
by: Avogaro, Niccolo, et al.
Published: (2026)
by: Avogaro, Niccolo, et al.
Published: (2026)
VP Lab: a PEFT-Enabled Visual Prompting Laboratory for Semantic Segmentation
by: Avogaro, Niccolo, et al.
Published: (2025)
by: Avogaro, Niccolo, et al.
Published: (2025)
Prompting-based Synthetic Data Generation for Few-Shot Question Answering
by: Schmidt, Maximilian, et al.
Published: (2024)
by: Schmidt, Maximilian, et al.
Published: (2024)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models
by: Farronato, Nicola, et al.
Published: (2026)
by: Farronato, Nicola, et al.
Published: (2026)
Combining Data Generation and Active Learning for Low-Resource Question Answering
by: Kimmich, Maximilian, et al.
Published: (2022)
by: Kimmich, Maximilian, et al.
Published: (2022)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
by: Cui, Shiyao, et al.
Published: (2025)
by: Cui, Shiyao, et al.
Published: (2025)
From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
by: Sheta, Hala, et al.
Published: (2025)
by: Sheta, Hala, et al.
Published: (2025)
Evolutionary Prompt Optimization Discovers Emergent Multimodal Reasoning Strategies in Vision-Language Models
by: Bharthulwar, Sid, et al.
Published: (2025)
by: Bharthulwar, Sid, et al.
Published: (2025)
Are complicated loss functions necessary for teaching LLMs to reason?
by: Carrino, Gabriele, et al.
Published: (2026)
by: Carrino, Gabriele, et al.
Published: (2026)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
by: Sood, Ekta, et al.
Published: (2021)
by: Sood, Ekta, et al.
Published: (2021)
PoseLess: Depth-Free Vision-to-Joint Control via Direct Image Mapping with VLM
by: Dao, Alan, et al.
Published: (2025)
by: Dao, Alan, et al.
Published: (2025)
EnsemHalDet: Robust VLM Hallucination Detection via Ensemble of Internal State Detectors
by: Miyazato, Ryuhei, et al.
Published: (2026)
by: Miyazato, Ryuhei, et al.
Published: (2026)
Voila-A: Aligning Vision-Language Models with User's Gaze Attention
by: Yan, Kun, et al.
Published: (2023)
by: Yan, Kun, et al.
Published: (2023)
Progressive Multimodal Reasoning via Active Retrieval
by: Dong, Guanting, et al.
Published: (2024)
by: Dong, Guanting, et al.
Published: (2024)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM
by: Wu, Chengyue, et al.
Published: (2026)
by: Wu, Chengyue, et al.
Published: (2026)
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
by: Jiang, Ziyan, et al.
Published: (2024)
by: Jiang, Ziyan, et al.
Published: (2024)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
by: Tian, Changyuan, et al.
Published: (2026)
by: Tian, Changyuan, et al.
Published: (2026)
MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models
by: Ahmed, Seif, et al.
Published: (2025)
by: Ahmed, Seif, et al.
Published: (2025)
PersonaVLM: Long-Term Personalized Multimodal LLMs
by: Nie, Chang, et al.
Published: (2026)
by: Nie, Chang, et al.
Published: (2026)
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
by: Hu, Zhe, et al.
Published: (2025)
by: Hu, Zhe, et al.
Published: (2025)
Detecting Hallucinations in Large Language Models via Internal Attention Divergence Signals
by: van Dijk, Gijs
Published: (2026)
by: van Dijk, Gijs
Published: (2026)
Controlling Reading Ease with Gaze-Guided Text Generation
by: Säuberli, Andreas, et al.
Published: (2026)
by: Säuberli, Andreas, et al.
Published: (2026)
Hybrid Decision Making via Conformal VLM-generated Guidance
by: Banerjee, Debodeep, et al.
Published: (2026)
by: Banerjee, Debodeep, et al.
Published: (2026)
MELO: An Evaluation Benchmark for Multilingual Entity Linking of Occupations
by: Retyk, Federico, et al.
Published: (2024)
by: Retyk, Federico, et al.
Published: (2024)
RoRA-VLM: Robust Retrieval-Augmented Vision Language Models
by: Qi, Jingyuan, et al.
Published: (2024)
by: Qi, Jingyuan, et al.
Published: (2024)
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
by: Qu, Kevin, et al.
Published: (2026)
by: Qu, Kevin, et al.
Published: (2026)
GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder
by: Cho, Seunghyuk, et al.
Published: (2025)
by: Cho, Seunghyuk, et al.
Published: (2025)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
by: Liu, Wenjie, et al.
Published: (2026)
by: Liu, Wenjie, et al.
Published: (2026)
Coherent Multimodal Reasoning with Iterative Self-Evaluation for Vision-Language Models
by: Luo, Wenjie, et al.
Published: (2025)
by: Luo, Wenjie, et al.
Published: (2025)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow Networks
by: Kang, Haoqiang, et al.
Published: (2025)
by: Kang, Haoqiang, et al.
Published: (2025)
Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
by: Song, Yusheng, et al.
Published: (2025)
by: Song, Yusheng, et al.
Published: (2025)
Similar Items
-
Eliciting Reasoning in Language Models with Cognitive Tools
by: Ebouky, Brown, et al.
Published: (2025) -
Show or Tell? Effectively prompting Vision-Language Models for semantic segmentation
by: Avogaro, Niccolo, et al.
Published: (2025) -
Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
by: Ebouky, Brown, et al.
Published: (2025) -
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
by: Mathew, Athul M., et al.
Published: (2025) -
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
by: Avogaro, Niccolo, et al.
Published: (2026)