Voila-A: Aligning Vision-Language Models with User's Gaze Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Kun, Ji, Lei, Wang, Zeyu, Wang, Yuntao, Duan, Nan, Ma, Shuai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
von: Ebouky, Brown, et al.
Veröffentlicht: (2026)
von: Ebouky, Brown, et al.
Veröffentlicht: (2026)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
von: Zhang, Zory, et al.
Veröffentlicht: (2025)
von: Zhang, Zory, et al.
Veröffentlicht: (2025)
Cost-effective Instruction Learning for Pathology Vision and Language Analysis
von: Chen, Kaitao, et al.
Veröffentlicht: (2024)
von: Chen, Kaitao, et al.
Veröffentlicht: (2024)
Enhancing Human-Computer Interaction in Chest X-ray Analysis using Vision and Language Model with Eye Gaze Patterns
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
Probing and Inducing Combinational Creativity in Vision-Language Models
von: Peng, Yongqian, et al.
Veröffentlicht: (2025)
von: Peng, Yongqian, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
von: An, Wenbin, et al.
Veröffentlicht: (2024)
von: An, Wenbin, et al.
Veröffentlicht: (2024)
Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update
von: Li, Qing, et al.
Veröffentlicht: (2025)
von: Li, Qing, et al.
Veröffentlicht: (2025)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models
von: Fu, Shuai, et al.
Veröffentlicht: (2024)
von: Fu, Shuai, et al.
Veröffentlicht: (2024)
Aesthetic Assessment of Chinese Handwritings Based on Vision Language Models
von: Zheng, Chen, et al.
Veröffentlicht: (2026)
von: Zheng, Chen, et al.
Veröffentlicht: (2026)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2025)
von: Woo, Sangmin, et al.
Veröffentlicht: (2025)
CANVAS: A Benchmark for Vision-Language Models on Tool-Based User Interface Design
von: Jeong, Daeheon, et al.
Veröffentlicht: (2025)
von: Jeong, Daeheon, et al.
Veröffentlicht: (2025)
MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization
von: Zhu, Kangyu, et al.
Veröffentlicht: (2024)
von: Zhu, Kangyu, et al.
Veröffentlicht: (2024)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
von: Zhong, Weihong, et al.
Veröffentlicht: (2024)
von: Zhong, Weihong, et al.
Veröffentlicht: (2024)
Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language Models
von: Liu, Xinyang, et al.
Veröffentlicht: (2023)
von: Liu, Xinyang, et al.
Veröffentlicht: (2023)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
An Examination of the Compositionality of Large Generative Vision-Language Models
von: Ma, Teli, et al.
Veröffentlicht: (2023)
von: Ma, Teli, et al.
Veröffentlicht: (2023)
Are Large Vision Language Models Good Game Players?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
Resolving Ambiguity in Gaze-Facilitated Visual Assistant Interaction Paradigm
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
von: Fan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Fan, Zhiwen, et al.
Veröffentlicht: (2025)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
von: Wang, Peng, et al.
Veröffentlicht: (2024)
von: Wang, Peng, et al.
Veröffentlicht: (2024)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
ChartInsights: Evaluating Multimodal Large Language Models for Low-Level Chart Question Answering
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
von: Li, Hongxing, et al.
Veröffentlicht: (2025)
von: Li, Hongxing, et al.
Veröffentlicht: (2025)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
von: Ye, Zekai, et al.
Veröffentlicht: (2025)
von: Ye, Zekai, et al.
Veröffentlicht: (2025)
Vision Token Reduction via Attention-Driven Self-Compression for Efficient Multimodal Large Language Models
von: Deniz, Omer Faruk, et al.
Veröffentlicht: (2026)
von: Deniz, Omer Faruk, et al.
Veröffentlicht: (2026)
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
von: Jiang, Songtao, et al.
Veröffentlicht: (2025)
von: Jiang, Songtao, et al.
Veröffentlicht: (2025)
Multi-Object Hallucination in Vision-Language Models
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents
von: Ma, Tianyi, et al.
Veröffentlicht: (2025)
von: Ma, Tianyi, et al.
Veröffentlicht: (2025)
CANAMRF: An Attention-Based Model for Multimodal Depression Detection
von: Wei, Yuntao, et al.
Veröffentlicht: (2024)
von: Wei, Yuntao, et al.
Veröffentlicht: (2024)
Can Vision-Language Models Solve Visual Math Equations?
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
von: Wu, Jiulong, et al.
Veröffentlicht: (2025)
von: Wu, Jiulong, et al.
Veröffentlicht: (2025)
OmniGen2: Towards Instruction-Aligned Multimodal Generation
von: Wu, Chenyuan, et al.
Veröffentlicht: (2025)
von: Wu, Chenyuan, et al.
Veröffentlicht: (2025)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
Large Vision-Language Models Get Lost in Attention
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
von: Ebouky, Brown, et al.
Veröffentlicht: (2026) -
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
von: Li, Bin, et al.
Veröffentlicht: (2025) -
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
von: Zhang, Zory, et al.
Veröffentlicht: (2025) -
Cost-effective Instruction Learning for Pathology Vision and Language Analysis
von: Chen, Kaitao, et al.
Veröffentlicht: (2024) -
Enhancing Human-Computer Interaction in Chest X-ray Analysis using Vision and Language Model with Eye Gaze Patterns
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)