When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Yakun, Cui, Qi, Xingqun, Geng, TianTian, Zhang, Yuyao, Han, Sirui, Guo, Yike |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
di: Wang, Zhiqiang, et al.
Pubblicazione: (2026)
di: Wang, Zhiqiang, et al.
Pubblicazione: (2026)
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
di: Shu, Yan, et al.
Pubblicazione: (2025)
di: Shu, Yan, et al.
Pubblicazione: (2025)
Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
di: Zhang, Chengsheng, et al.
Pubblicazione: (2026)
di: Zhang, Chengsheng, et al.
Pubblicazione: (2026)
From Clouds to Hallucinations: Atmospheric Retrieval Hijacking in Remote Sensing Vision-Language RAG
di: Han, Jiaju, et al.
Pubblicazione: (2026)
di: Han, Jiaju, et al.
Pubblicazione: (2026)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
di: Wu, Yike, et al.
Pubblicazione: (2025)
di: Wu, Yike, et al.
Pubblicazione: (2025)
Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation
di: Chen, Beitao, et al.
Pubblicazione: (2025)
di: Chen, Beitao, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
di: Lu, Yifan, et al.
Pubblicazione: (2025)
di: Lu, Yifan, et al.
Pubblicazione: (2025)
Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection
di: Yakun, Cui, et al.
Pubblicazione: (2025)
di: Yakun, Cui, et al.
Pubblicazione: (2025)
Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models
di: Hu, Nanxing, et al.
Pubblicazione: (2025)
di: Hu, Nanxing, et al.
Pubblicazione: (2025)
Emergence of Text Readability in Vision Language Models
di: Park, Jaeyoo, et al.
Pubblicazione: (2025)
di: Park, Jaeyoo, et al.
Pubblicazione: (2025)
Mitigating Image Captioning Hallucinations in Vision-Language Models
di: Zhao, Fei, et al.
Pubblicazione: (2025)
di: Zhao, Fei, et al.
Pubblicazione: (2025)
Tag2Text: Guiding Vision-Language Model via Image Tagging
di: Huang, Xinyu, et al.
Pubblicazione: (2023)
di: Huang, Xinyu, et al.
Pubblicazione: (2023)
OTR: Synthesizing Overlay Text Dataset for Text Removal
di: Zdenek, Jan, et al.
Pubblicazione: (2025)
di: Zdenek, Jan, et al.
Pubblicazione: (2025)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
di: Zhou, Yiyang, et al.
Pubblicazione: (2023)
di: Zhou, Yiyang, et al.
Pubblicazione: (2023)
HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation
di: Wang, Haoyu, et al.
Pubblicazione: (2026)
di: Wang, Haoyu, et al.
Pubblicazione: (2026)
Exploring Causes and Mitigation of Hallucinations in Large Vision Language Models
di: Sun, Yaqi, et al.
Pubblicazione: (2025)
di: Sun, Yaqi, et al.
Pubblicazione: (2025)
YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models
di: Chen, Ting, et al.
Pubblicazione: (2026)
di: Chen, Ting, et al.
Pubblicazione: (2026)
MedHEval: Benchmarking Hallucinations and Mitigation Strategies in Medical Large Vision-Language Models
di: Chang, Aofei, et al.
Pubblicazione: (2025)
di: Chang, Aofei, et al.
Pubblicazione: (2025)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
di: Park, Dongmin, et al.
Pubblicazione: (2024)
di: Park, Dongmin, et al.
Pubblicazione: (2024)
Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
di: Lyu, Xinyu, et al.
Pubblicazione: (2024)
di: Lyu, Xinyu, et al.
Pubblicazione: (2024)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
di: An, Wenbin, et al.
Pubblicazione: (2024)
di: An, Wenbin, et al.
Pubblicazione: (2024)
ReCo: Reminder Composition Mitigates Hallucinations in Vision-Language Models
di: Chytas, Sotirios Panagiotis, et al.
Pubblicazione: (2025)
di: Chytas, Sotirios Panagiotis, et al.
Pubblicazione: (2025)
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
di: Zhang, Jinrui, et al.
Pubblicazione: (2024)
di: Zhang, Jinrui, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
di: Zhu, Xingyu, et al.
Pubblicazione: (2026)
di: Zhu, Xingyu, et al.
Pubblicazione: (2026)
Mitigating Multilingual Hallucination in Large Vision-Language Models
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
di: Yang, Zhihe, et al.
Pubblicazione: (2025)
di: Yang, Zhihe, et al.
Pubblicazione: (2025)
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
di: Wang, Weihang, et al.
Pubblicazione: (2025)
di: Wang, Weihang, et al.
Pubblicazione: (2025)
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
di: Xu, Juangui, et al.
Pubblicazione: (2025)
di: Xu, Juangui, et al.
Pubblicazione: (2025)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
di: Maeda, Koki, et al.
Pubblicazione: (2026)
di: Maeda, Koki, et al.
Pubblicazione: (2026)
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
di: Tian, Xinyu, et al.
Pubblicazione: (2025)
di: Tian, Xinyu, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Causal Route Gating
di: Cheng, Zhe, et al.
Pubblicazione: (2026)
di: Cheng, Zhe, et al.
Pubblicazione: (2026)
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
di: Jiang, Nick, et al.
Pubblicazione: (2024)
di: Jiang, Nick, et al.
Pubblicazione: (2024)
VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking
di: Fu, Jiyuan, et al.
Pubblicazione: (2026)
di: Fu, Jiyuan, et al.
Pubblicazione: (2026)
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs
di: Choong, Wey Yeh, et al.
Pubblicazione: (2024)
di: Choong, Wey Yeh, et al.
Pubblicazione: (2024)
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
di: Yan, Qiao, et al.
Pubblicazione: (2025)
di: Yan, Qiao, et al.
Pubblicazione: (2025)
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
di: Liu, Yufang, et al.
Pubblicazione: (2024)
di: Liu, Yufang, et al.
Pubblicazione: (2024)
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
di: Zhang, Yuanhong, et al.
Pubblicazione: (2026)
di: Zhang, Yuanhong, et al.
Pubblicazione: (2026)
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
di: Chang, Yue, et al.
Pubblicazione: (2024)
di: Chang, Yue, et al.
Pubblicazione: (2024)
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
di: Li, Chenxi, et al.
Pubblicazione: (2025)
di: Li, Chenxi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
di: Wang, Zhiqiang, et al.
Pubblicazione: (2026) -
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
di: Shu, Yan, et al.
Pubblicazione: (2025) -
Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
di: Zhang, Chengsheng, et al.
Pubblicazione: (2026) -
From Clouds to Hallucinations: Atmospheric Retrieval Hijacking in Remote Sensing Vision-Language RAG
di: Han, Jiaju, et al.
Pubblicazione: (2026) -
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
di: Wu, Yike, et al.
Pubblicazione: (2025)