SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
Fuente:
arXiv
Guardado en:
| Autores principales: | Shukla, Tripti, Kira, Zsolt |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
por: Tong, Bingkui, et al.
Publicado: (2025)
por: Tong, Bingkui, et al.
Publicado: (2025)
SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
por: Ye, Wenqian, et al.
Publicado: (2025)
por: Ye, Wenqian, et al.
Publicado: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
por: Kumar, Akash, et al.
Publicado: (2025)
por: Kumar, Akash, et al.
Publicado: (2025)
VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs
por: Kolli, Govinda, et al.
Publicado: (2026)
por: Kolli, Govinda, et al.
Publicado: (2026)
Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding
por: Li, Jinlin, et al.
Publicado: (2025)
por: Li, Jinlin, et al.
Publicado: (2025)
Energy-Guided Decoding for Object Hallucination Mitigation
por: Liu, Xixi, et al.
Publicado: (2025)
por: Liu, Xixi, et al.
Publicado: (2025)
Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding
por: Xu, Zhongxing, et al.
Publicado: (2026)
por: Xu, Zhongxing, et al.
Publicado: (2026)
SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding
por: Li, Zhaoxu, et al.
Publicado: (2026)
por: Li, Zhaoxu, et al.
Publicado: (2026)
Test-time Conditional Text-to-Image Synthesis Using Diffusion Models
por: Shukla, Tripti, et al.
Publicado: (2024)
por: Shukla, Tripti, et al.
Publicado: (2024)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
por: Park, Yeji, et al.
Publicado: (2024)
por: Park, Yeji, et al.
Publicado: (2024)
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
por: Ma, Ji, et al.
Publicado: (2026)
por: Ma, Ji, et al.
Publicado: (2026)
Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting
por: Wu, Jiarui, et al.
Publicado: (2025)
por: Wu, Jiarui, et al.
Publicado: (2025)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
por: Feng, Wenfeng, et al.
Publicado: (2025)
por: Feng, Wenfeng, et al.
Publicado: (2025)
Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding
por: Liang, Xiaoyu, et al.
Publicado: (2024)
por: Liang, Xiaoyu, et al.
Publicado: (2024)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
por: Chen, Boqi, et al.
Publicado: (2026)
por: Chen, Boqi, et al.
Publicado: (2026)
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
por: Mao, Jiawei, et al.
Publicado: (2026)
por: Mao, Jiawei, et al.
Publicado: (2026)
On the Nature of Attention Sink that Shapes Decoding Strategy in Omni-LLMs
por: Yoo, Suho, et al.
Publicado: (2026)
por: Yoo, Suho, et al.
Publicado: (2026)
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
por: Chen, Xinrong, et al.
Publicado: (2026)
por: Chen, Xinrong, et al.
Publicado: (2026)
Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation
por: Wu, Yebo, et al.
Publicado: (2026)
por: Wu, Yebo, et al.
Publicado: (2026)
AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
por: Jung, Chaeyoung, et al.
Publicado: (2025)
por: Jung, Chaeyoung, et al.
Publicado: (2025)
SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
por: Park, Woohyeon, et al.
Publicado: (2025)
por: Park, Woohyeon, et al.
Publicado: (2025)
Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
por: Yang, Tiancheng, et al.
Publicado: (2025)
por: Yang, Tiancheng, et al.
Publicado: (2025)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
por: He, Zhentao, et al.
Publicado: (2025)
por: He, Zhentao, et al.
Publicado: (2025)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
por: Cai, Jianfeng, et al.
Publicado: (2025)
por: Cai, Jianfeng, et al.
Publicado: (2025)
Mitigating Multimodal Hallucination via Phase-wise Self-reward
por: Zhang, Yu, et al.
Publicado: (2026)
por: Zhang, Yu, et al.
Publicado: (2026)
Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection
por: Wang, Shan, et al.
Publicado: (2025)
por: Wang, Shan, et al.
Publicado: (2025)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
por: Whitehead, Spencer, et al.
Publicado: (2024)
por: Whitehead, Spencer, et al.
Publicado: (2024)
YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models
por: Chen, Ting, et al.
Publicado: (2026)
por: Chen, Ting, et al.
Publicado: (2026)
MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue
por: Jiang, Yue, et al.
Publicado: (2026)
por: Jiang, Yue, et al.
Publicado: (2026)
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
por: Liu, Yi, et al.
Publicado: (2026)
por: Liu, Yi, et al.
Publicado: (2026)
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
por: Li, Jiale, et al.
Publicado: (2025)
por: Li, Jiale, et al.
Publicado: (2025)
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
por: Gao, Hongcheng, et al.
Publicado: (2025)
por: Gao, Hongcheng, et al.
Publicado: (2025)
GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations
por: Chen, Boyuan, et al.
Publicado: (2026)
por: Chen, Boyuan, et al.
Publicado: (2026)
SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
por: Tong, Yujia, et al.
Publicado: (2026)
por: Tong, Yujia, et al.
Publicado: (2026)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
por: Tang, Feilong, et al.
Publicado: (2025)
por: Tang, Feilong, et al.
Publicado: (2025)
Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding
por: Cho, Yeongjae, et al.
Publicado: (2025)
por: Cho, Yeongjae, et al.
Publicado: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
por: Wang, Chao, et al.
Publicado: (2025)
por: Wang, Chao, et al.
Publicado: (2025)
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
por: Zhang, Ce, et al.
Publicado: (2025)
por: Zhang, Ce, et al.
Publicado: (2025)
Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
por: Ma, Ruiqi, et al.
Publicado: (2025)
por: Ma, Ruiqi, et al.
Publicado: (2025)
HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
por: Yuan, Fan, et al.
Publicado: (2024)
por: Yuan, Fan, et al.
Publicado: (2024)
Ejemplares similares
-
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
por: Tong, Bingkui, et al.
Publicado: (2025) -
SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
por: Ye, Wenqian, et al.
Publicado: (2025) -
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
por: Kumar, Akash, et al.
Publicado: (2025) -
VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs
por: Kolli, Govinda, et al.
Publicado: (2026) -
Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding
por: Li, Jinlin, et al.
Publicado: (2025)