Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Zhaochen, Wang, Yiwei, Cai, Yujun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
por: Wang, Zhaochen, et al.
Publicado: (2025)
por: Wang, Zhaochen, et al.
Publicado: (2025)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
por: Wu, Yike, et al.
Publicado: (2025)
por: Wu, Yike, et al.
Publicado: (2025)
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
por: Ge, Haonan, et al.
Publicado: (2025)
por: Ge, Haonan, et al.
Publicado: (2025)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
por: Wu, Hang, et al.
Publicado: (2026)
por: Wu, Hang, et al.
Publicado: (2026)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
por: Wang, Xintong, et al.
Publicado: (2024)
por: Wang, Xintong, et al.
Publicado: (2024)
Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs
por: Zhang, Kejia, et al.
Publicado: (2025)
por: Zhang, Kejia, et al.
Publicado: (2025)
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
por: Yan, Qiao, et al.
Publicado: (2025)
por: Yan, Qiao, et al.
Publicado: (2025)
EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models
por: Villa, Andrés, et al.
Publicado: (2025)
por: Villa, Andrés, et al.
Publicado: (2025)
Prescribing the Right Remedy: Mitigating Hallucinations in Large Vision-Language Models via Targeted Instruction Tuning
por: Hu, Rui, et al.
Publicado: (2024)
por: Hu, Rui, et al.
Publicado: (2024)
From Clouds to Hallucinations: Atmospheric Retrieval Hijacking in Remote Sensing Vision-Language RAG
por: Han, Jiaju, et al.
Publicado: (2026)
por: Han, Jiaju, et al.
Publicado: (2026)
TPC: Cross-Temporal Prediction Connection for Vision-Language Model Hallucination Reduction
por: Wang, Chao, et al.
Publicado: (2025)
por: Wang, Chao, et al.
Publicado: (2025)
Hierarchical, Interpretable, Label-Free Concept Bottleneck Model
por: Xie, Haodong, et al.
Publicado: (2026)
por: Xie, Haodong, et al.
Publicado: (2026)
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
por: Zhang, Yuanhong, et al.
Publicado: (2026)
por: Zhang, Yuanhong, et al.
Publicado: (2026)
Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
por: Huo, Fushuo, et al.
Publicado: (2024)
por: Huo, Fushuo, et al.
Publicado: (2024)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
por: Woo, Sangmin, et al.
Publicado: (2025)
por: Woo, Sangmin, et al.
Publicado: (2025)
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
por: Sun, Bowen, et al.
Publicado: (2025)
por: Sun, Bowen, et al.
Publicado: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
por: Ge, Haonan, et al.
Publicado: (2025)
por: Ge, Haonan, et al.
Publicado: (2025)
SDCD: Structure-Disrupted Contrastive Decoding for Mitigating Hallucinations in Large Vision-Language Models
por: Xia, Yuxuan, et al.
Publicado: (2026)
por: Xia, Yuxuan, et al.
Publicado: (2026)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
por: Seo, Hoigi, et al.
Publicado: (2025)
por: Seo, Hoigi, et al.
Publicado: (2025)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
por: Lee, Yi-Lun, et al.
Publicado: (2024)
por: Lee, Yi-Lun, et al.
Publicado: (2024)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
por: Zhang, Yudong, et al.
Publicado: (2024)
por: Zhang, Yudong, et al.
Publicado: (2024)
Review of Hallucination Understanding in Large Language and Vision Models
por: Ho, Zhengyi, et al.
Publicado: (2025)
por: Ho, Zhengyi, et al.
Publicado: (2025)
DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models
por: Wang, JiYang, et al.
Publicado: (2026)
por: Wang, JiYang, et al.
Publicado: (2026)
VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information Bottleneck
por: Zhang, Feiran, et al.
Publicado: (2026)
por: Zhang, Feiran, et al.
Publicado: (2026)
INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling
por: Dong, Xin, et al.
Publicado: (2025)
por: Dong, Xin, et al.
Publicado: (2025)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
por: Fu, Honghao, et al.
Publicado: (2026)
por: Fu, Honghao, et al.
Publicado: (2026)
Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models
por: Hu, Nanxing, et al.
Publicado: (2025)
por: Hu, Nanxing, et al.
Publicado: (2025)
Revealing Multi-View Hallucination in Large Vision-Language Models
por: Park, Wooje, et al.
Publicado: (2026)
por: Park, Wooje, et al.
Publicado: (2026)
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
por: Li, Zhuowei, et al.
Publicado: (2025)
por: Li, Zhuowei, et al.
Publicado: (2025)
V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
por: Sun, Nan, et al.
Publicado: (2025)
por: Sun, Nan, et al.
Publicado: (2025)
Multi-Object Hallucination in Vision-Language Models
por: Chen, Xuweiyi, et al.
Publicado: (2024)
por: Chen, Xuweiyi, et al.
Publicado: (2024)
MURE: Hierarchical Multi-Resolution Encoding via Vision-Language Models for Visual Document Retrieval
por: Zhu, Fengbin, et al.
Publicado: (2026)
por: Zhu, Fengbin, et al.
Publicado: (2026)
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
por: Lim, Sohwi, et al.
Publicado: (2026)
por: Lim, Sohwi, et al.
Publicado: (2026)
Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
por: Zhang, Chengsheng, et al.
Publicado: (2026)
por: Zhang, Chengsheng, et al.
Publicado: (2026)
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
por: Liu, Yufang, et al.
Publicado: (2024)
por: Liu, Yufang, et al.
Publicado: (2024)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
por: Seth, Ashish, et al.
Publicado: (2024)
por: Seth, Ashish, et al.
Publicado: (2024)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
por: Zhang, Jieyu, et al.
Publicado: (2024)
por: Zhang, Jieyu, et al.
Publicado: (2024)
Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
por: Chae, Hyunsik, et al.
Publicado: (2025)
por: Chae, Hyunsik, et al.
Publicado: (2025)
Mitigating Multilingual Hallucination in Large Vision-Language Models
por: Qu, Xiaoye, et al.
Publicado: (2024)
por: Qu, Xiaoye, et al.
Publicado: (2024)
Benchmarking Deflection and Hallucination in Large Vision-Language Models
por: Moratelli, Nicholas, et al.
Publicado: (2026)
por: Moratelli, Nicholas, et al.
Publicado: (2026)
Ejemplares similares
-
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
por: Wang, Zhaochen, et al.
Publicado: (2025) -
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
por: Wu, Yike, et al.
Publicado: (2025) -
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
por: Ge, Haonan, et al.
Publicado: (2025) -
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
por: Wu, Hang, et al.
Publicado: (2026) -
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
por: Wang, Xintong, et al.
Publicado: (2024)