Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Yeongjae, Kim, Keonwoo, Hwang, Taebaek, Cho, Sungzoon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds
by: Kim, Keonwoo, et al.
Published: (2025)
by: Kim, Keonwoo, et al.
Published: (2025)
Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays
by: Cho, Yeongjae, et al.
Published: (2024)
by: Cho, Yeongjae, et al.
Published: (2024)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
by: Kim, Hyunjong, et al.
Published: (2025)
by: Kim, Hyunjong, et al.
Published: (2025)
An Interpretable Local Editing Model for Counterfactual Medical Image Generation
by: Min, Hyungi, et al.
Published: (2026)
by: Min, Hyungi, et al.
Published: (2026)
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
by: Cho, Beomsik, et al.
Published: (2025)
by: Cho, Beomsik, et al.
Published: (2025)
Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding
by: Li, Jinlin, et al.
Published: (2025)
by: Li, Jinlin, et al.
Published: (2025)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Energy-Guided Decoding for Object Hallucination Mitigation
by: Liu, Xixi, et al.
Published: (2025)
by: Liu, Xixi, et al.
Published: (2025)
Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation
by: Wu, Yebo, et al.
Published: (2026)
by: Wu, Yebo, et al.
Published: (2026)
A Revisit to the Decoder for Camouflaged Object Detection
by: Ko, Seung Woo, et al.
Published: (2025)
by: Ko, Seung Woo, et al.
Published: (2025)
Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models
by: Magid, Salma Abdel, et al.
Published: (2024)
by: Magid, Salma Abdel, et al.
Published: (2024)
Generalizing Visual Question Answering from Synthetic to Human-Written Questions via a Chain of QA with a Large Language Model
by: Kim, Taehee, et al.
Published: (2024)
by: Kim, Taehee, et al.
Published: (2024)
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
by: Tong, Bingkui, et al.
Published: (2025)
by: Tong, Bingkui, et al.
Published: (2025)
Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
by: Yang, Tiancheng, et al.
Published: (2025)
by: Yang, Tiancheng, et al.
Published: (2025)
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
by: Shukla, Tripti, et al.
Published: (2026)
by: Shukla, Tripti, et al.
Published: (2026)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
by: Oh, Hongseok, et al.
Published: (2025)
by: Oh, Hongseok, et al.
Published: (2025)
SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
by: Park, Woohyeon, et al.
Published: (2025)
by: Park, Woohyeon, et al.
Published: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
by: Min, Kyungmin, et al.
Published: (2024)
by: Min, Kyungmin, et al.
Published: (2024)
KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
by: Hwang, Taebaek, et al.
Published: (2025)
by: Hwang, Taebaek, et al.
Published: (2025)
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025)
by: Kala, Divy, et al.
Published: (2025)
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
by: Song, Chull Hwan, et al.
Published: (2024)
by: Song, Chull Hwan, et al.
Published: (2024)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
by: Tang, Feilong, et al.
Published: (2025)
by: Tang, Feilong, et al.
Published: (2025)
VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs
by: Kolli, Govinda, et al.
Published: (2026)
by: Kolli, Govinda, et al.
Published: (2026)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
by: Zhou, Guanyu, et al.
Published: (2024)
by: Zhou, Guanyu, et al.
Published: (2024)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
by: Chen, Boqi, et al.
Published: (2026)
by: Chen, Boqi, et al.
Published: (2026)
Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal Hallucination
by: Shi, Ziqiang, et al.
Published: (2026)
by: Shi, Ziqiang, et al.
Published: (2026)
Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs
by: Lin, Chenchen, et al.
Published: (2026)
by: Lin, Chenchen, et al.
Published: (2026)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
by: Li, Qiming, et al.
Published: (2026)
by: Li, Qiming, et al.
Published: (2026)
Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding
by: Liang, Xiaoyu, et al.
Published: (2024)
by: Liang, Xiaoyu, et al.
Published: (2024)
Anatomical Region-Guided Contrastive Decoding: A Plug-and-Play Strategy for Mitigating Hallucinations in Medical VLMs
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
Mitigating Object Hallucination via Concentric Causal Attention
by: Xing, Yun, et al.
Published: (2024)
by: Xing, Yun, et al.
Published: (2024)
Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision
by: Lee, Seongyun, et al.
Published: (2023)
by: Lee, Seongyun, et al.
Published: (2023)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
by: Park, Yeji, et al.
Published: (2024)
by: Park, Yeji, et al.
Published: (2024)
Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation
by: Chen, Beitao, et al.
Published: (2025)
by: Chen, Beitao, et al.
Published: (2025)
Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
by: Yang, Chengxu, et al.
Published: (2026)
by: Yang, Chengxu, et al.
Published: (2026)
Robust Multimodal 3D Object Detection via Modality-Agnostic Decoding and Proximity-based Modality Ensemble
by: Cha, Juhan, et al.
Published: (2024)
by: Cha, Juhan, et al.
Published: (2024)
Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification
by: Sun, Han, et al.
Published: (2026)
by: Sun, Han, et al.
Published: (2026)
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
by: Zou, Xin, et al.
Published: (2024)
by: Zou, Xin, et al.
Published: (2024)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026)
by: Song, Junha, et al.
Published: (2026)
Similar Items
-
CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds
by: Kim, Keonwoo, et al.
Published: (2025) -
Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays
by: Cho, Yeongjae, et al.
Published: (2024) -
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
by: Kim, Hyunjong, et al.
Published: (2025) -
An Interpretable Local Editing Model for Counterfactual Medical Image Generation
by: Min, Hyungi, et al.
Published: (2026) -
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
by: Cho, Beomsik, et al.
Published: (2025)