Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
Fuente:
arXiv
Guardado en:
| Autores principales: | Qi, Zheng, Shang, Chao, Spiliopoulou, Evangelia, Pappas, Nikolaos |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance
por: Cao, Jinjin, et al.
Publicado: (2025)
por: Cao, Jinjin, et al.
Publicado: (2025)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
por: He, Lehan, et al.
Publicado: (2024)
por: He, Lehan, et al.
Publicado: (2024)
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
por: Li, Jiaming, et al.
Publicado: (2025)
por: Li, Jiaming, et al.
Publicado: (2025)
Leveraging Multi-Modal Saliency and Fusion for Gaze Target Detection
por: Mathew, Athul M., et al.
Publicado: (2025)
por: Mathew, Athul M., et al.
Publicado: (2025)
InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration
por: Yang, Zhongyu, et al.
Publicado: (2025)
por: Yang, Zhongyu, et al.
Publicado: (2025)
Ocular Authentication: Fusion of Gaze and Periocular Modalities
por: Lohr, Dillon, et al.
Publicado: (2025)
por: Lohr, Dillon, et al.
Publicado: (2025)
Mitigating Diffusion Model Hallucinations with Dynamic Guidance
por: Triaridis, Kostas, et al.
Publicado: (2025)
por: Triaridis, Kostas, et al.
Publicado: (2025)
Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
por: Zheng, Haohan, et al.
Publicado: (2025)
por: Zheng, Haohan, et al.
Publicado: (2025)
Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models
por: Wang, Hengfei, et al.
Publicado: (2026)
por: Wang, Hengfei, et al.
Publicado: (2026)
GazeShift: Unsupervised Gaze Estimation and Dataset for VR
por: Shapira, Gil, et al.
Publicado: (2026)
por: Shapira, Gil, et al.
Publicado: (2026)
Gaze Label Alignment: Alleviating Domain Shift for Gaze Estimation
por: Zeng, Guanzhong, et al.
Publicado: (2024)
por: Zeng, Guanzhong, et al.
Publicado: (2024)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
por: Bae, Kyungho, et al.
Publicado: (2025)
por: Bae, Kyungho, et al.
Publicado: (2025)
VLM-UQBench: A Benchmark for Modality-Specific and Cross-Modality Uncertainties in Vision Language Models
por: Wang, Chenyu, et al.
Publicado: (2026)
por: Wang, Chenyu, et al.
Publicado: (2026)
INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling
por: Dong, Xin, et al.
Publicado: (2025)
por: Dong, Xin, et al.
Publicado: (2025)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
por: Pani, Anupam, et al.
Publicado: (2025)
por: Pani, Anupam, et al.
Publicado: (2025)
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
por: Mathew, Athul M., et al.
Publicado: (2025)
por: Mathew, Athul M., et al.
Publicado: (2025)
Integrating MedCLIP and Cross-Modal Fusion for Automatic Radiology Report Generation
por: Han, Qianhao, et al.
Publicado: (2024)
por: Han, Qianhao, et al.
Publicado: (2024)
Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement
por: Qin, Zhenxin, et al.
Publicado: (2026)
por: Qin, Zhenxin, et al.
Publicado: (2026)
BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
por: Singha, Mainak, et al.
Publicado: (2026)
por: Singha, Mainak, et al.
Publicado: (2026)
MMMamba: A Versatile Cross-Modal In Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement
por: Wang, Yingying, et al.
Publicado: (2025)
por: Wang, Yingying, et al.
Publicado: (2025)
CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition
por: Zheng, Jinzhi, et al.
Publicado: (2024)
por: Zheng, Jinzhi, et al.
Publicado: (2024)
Cross-Modal Guidance for Fast Diffusion-Based Computed Tomography
por: Efimov, Timofey, et al.
Publicado: (2026)
por: Efimov, Timofey, et al.
Publicado: (2026)
CrossFlowDG: Bridging the Modality Gap with Cross-modal Flow Matching for Domain Generalization
por: Kritikos, Antonios, et al.
Publicado: (2026)
por: Kritikos, Antonios, et al.
Publicado: (2026)
A Generalized Label Shift Perspective for Cross-Domain Gaze Estimation
por: Yang, Hao-Ran, et al.
Publicado: (2025)
por: Yang, Hao-Ran, et al.
Publicado: (2025)
Cross-Dataset Gaze Estimation by Evidential Inter-intra Fusion
por: Wang, Shijing, et al.
Publicado: (2024)
por: Wang, Shijing, et al.
Publicado: (2024)
Modality-Specific Enhancement and Complementary Fusion for Semi-Supervised Multi-Modal Brain Tumor Segmentation
por: Chung, Tien-Dat, et al.
Publicado: (2025)
por: Chung, Tien-Dat, et al.
Publicado: (2025)
Multi-Modal Gaze Following in Conversational Scenarios
por: Hou, Yuqi, et al.
Publicado: (2023)
por: Hou, Yuqi, et al.
Publicado: (2023)
Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection
por: Yang, Le, et al.
Publicado: (2024)
por: Yang, Le, et al.
Publicado: (2024)
CMIP-CIL: A Cross-Modal Benchmark for Image-Point Class Incremental Learning
por: Qi, Chao, et al.
Publicado: (2025)
por: Qi, Chao, et al.
Publicado: (2025)
Cross-Modal Mapping: Mitigating the Modality Gap for Few-Shot Image Classification
por: Yang, Xi, et al.
Publicado: (2024)
por: Yang, Xi, et al.
Publicado: (2024)
Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs
por: Lin, Chenchen, et al.
Publicado: (2026)
por: Lin, Chenchen, et al.
Publicado: (2026)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
por: Lu, Yifan, et al.
Publicado: (2025)
por: Lu, Yifan, et al.
Publicado: (2025)
Enriching Knowledge Distillation with Cross-Modal Teacher Fusion
por: Mansourian, Amir M., et al.
Publicado: (2025)
por: Mansourian, Amir M., et al.
Publicado: (2025)
GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection
por: Wang, Yu, et al.
Publicado: (2025)
por: Wang, Yu, et al.
Publicado: (2025)
Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs
por: Jo, Yujin, et al.
Publicado: (2026)
por: Jo, Yujin, et al.
Publicado: (2026)
Eye Gaze Tells You Where to Compute: Gaze-Driven Efficient VLMs
por: Chen, Qinyu, et al.
Publicado: (2025)
por: Chen, Qinyu, et al.
Publicado: (2025)
Interpreting and Mitigating Hallucination in MLLMs through Multi-agent Debate
por: Lin, Zheng, et al.
Publicado: (2024)
por: Lin, Zheng, et al.
Publicado: (2024)
Ultrasound Report Generation with Cross-Modality Feature Alignment via Unsupervised Guidance
por: Li, Jun, et al.
Publicado: (2024)
por: Li, Jun, et al.
Publicado: (2024)
MOS: Mitigating Optical-SAR Modality Gap for Cross-Modal Ship Re-Identification
por: Zhao, Yujian, et al.
Publicado: (2025)
por: Zhao, Yujian, et al.
Publicado: (2025)
CrossGaze: A Strong Method for 3D Gaze Estimation in the Wild
por: Cătrună, Andy, et al.
Publicado: (2024)
por: Cătrună, Andy, et al.
Publicado: (2024)
Ejemplares similares
-
When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance
por: Cao, Jinjin, et al.
Publicado: (2025) -
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
por: He, Lehan, et al.
Publicado: (2024) -
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
por: Li, Jiaming, et al.
Publicado: (2025) -
Leveraging Multi-Modal Saliency and Fusion for Gaze Target Detection
por: Mathew, Athul M., et al.
Publicado: (2025) -
InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration
por: Yang, Zhongyu, et al.
Publicado: (2025)