Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Chengxu, Yuan, Jingling, Hu, Chuang, Jiang, Jiawei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model
by: Liu, Haogeng, et al.
Published: (2024)
by: Liu, Haogeng, et al.
Published: (2024)
Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models
by: Liu, Zijian, et al.
Published: (2026)
by: Liu, Zijian, et al.
Published: (2026)
Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
by: Zhou, Haoran, et al.
Published: (2025)
by: Zhou, Haoran, et al.
Published: (2025)
Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation
by: Wu, Yebo, et al.
Published: (2026)
by: Wu, Yebo, et al.
Published: (2026)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
by: Li, Qiming, et al.
Published: (2026)
by: Li, Qiming, et al.
Published: (2026)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
by: Yang, Zhihe, et al.
Published: (2025)
by: Yang, Zhihe, et al.
Published: (2025)
Enhancing Quantization-Aware Training on Edge Devices via Relative Entropy Coreset Selection and Cascaded Layer Correction
by: Tong, Yujia, et al.
Published: (2025)
by: Tong, Yujia, et al.
Published: (2025)
AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Models
by: Qu, Zhen, et al.
Published: (2026)
by: Qu, Zhen, et al.
Published: (2026)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding
by: Liang, Xiaoyu, et al.
Published: (2024)
by: Liang, Xiaoyu, et al.
Published: (2024)
Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
by: Jiang, Zhangqi, et al.
Published: (2024)
by: Jiang, Zhangqi, et al.
Published: (2024)
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
by: Zhou, Guanyu, et al.
Published: (2024)
by: Zhou, Guanyu, et al.
Published: (2024)
Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning
by: Gong, Xuan, et al.
Published: (2026)
by: Gong, Xuan, et al.
Published: (2026)
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
by: Zou, Xin, et al.
Published: (2024)
by: Zou, Xin, et al.
Published: (2024)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
by: Park, Yeji, et al.
Published: (2024)
by: Park, Yeji, et al.
Published: (2024)
SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
by: Li, Zhaoxu, et al.
Published: (2025)
by: Li, Zhaoxu, et al.
Published: (2025)
V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
by: Sun, Nan, et al.
Published: (2025)
by: Sun, Nan, et al.
Published: (2025)
Language as an Anchor: Preserving Relative Visual Geometry for Domain Incremental Learning
by: Geng, Shuyi, et al.
Published: (2025)
by: Geng, Shuyi, et al.
Published: (2025)
Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
by: Yang, Tiancheng, et al.
Published: (2025)
by: Yang, Tiancheng, et al.
Published: (2025)
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
by: Shan, Jiquan, et al.
Published: (2025)
by: Shan, Jiquan, et al.
Published: (2025)
Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy
by: Xie, Yutong, et al.
Published: (2026)
by: Xie, Yutong, et al.
Published: (2026)
Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations
by: Yang, Chengxu, et al.
Published: (2025)
by: Yang, Chengxu, et al.
Published: (2025)
VaLiD: Mitigating the Hallucination of Large Vision Language Models by Visual Layer Fusion Contrastive Decoding
by: Wang, Jiaqi, et al.
Published: (2024)
by: Wang, Jiaqi, et al.
Published: (2024)
Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination
by: Yang, Dingchen, et al.
Published: (2024)
by: Yang, Dingchen, et al.
Published: (2024)
Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model
by: Zhuang, Jiedong, et al.
Published: (2026)
by: Zhuang, Jiedong, et al.
Published: (2026)
LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models
by: Guo, Zhihui, et al.
Published: (2025)
by: Guo, Zhihui, et al.
Published: (2025)
AnchorFlow: Editable SVG Reconstruction via Sparse Anchor Point Fields
by: Jiang, Mengnan, et al.
Published: (2026)
by: Jiang, Mengnan, et al.
Published: (2026)
Exposing Hallucinations To Suppress Them: VLMs Representation Editing With Generative Anchors
by: Shi, Youxu, et al.
Published: (2025)
by: Shi, Youxu, et al.
Published: (2025)
IKOD: Mitigating Visual Attention Degradation in Large Vision-Language Models
by: Yang, Jiabing, et al.
Published: (2025)
by: Yang, Jiabing, et al.
Published: (2025)
VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
by: Wang, Zhaozhi, et al.
Published: (2025)
by: Wang, Zhaozhi, et al.
Published: (2025)
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
by: You, Liangliang, et al.
Published: (2025)
by: You, Liangliang, et al.
Published: (2025)
Self-Correction Inside the Model: Leveraging Layer Attention to Mitigate Hallucinations in Large Vision Language Models
by: Fu, April
Published: (2026)
by: Fu, April
Published: (2026)
Anchor-based Robust Finetuning of Vision-Language Models
by: Han, Jinwei, et al.
Published: (2024)
by: Han, Jinwei, et al.
Published: (2024)
Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation
by: Nadeem, Numair, et al.
Published: (2025)
by: Nadeem, Numair, et al.
Published: (2025)
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
by: Deng, Haolin, et al.
Published: (2026)
by: Deng, Haolin, et al.
Published: (2026)
Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
by: Jung, Chaeyoung, et al.
Published: (2025)
by: Jung, Chaeyoung, et al.
Published: (2025)
VORD: Visual Ordinal Calibration for Mitigating Object Hallucinations in Large Vision-Language Models
by: Neo, Dexter, et al.
Published: (2024)
by: Neo, Dexter, et al.
Published: (2024)
Similar Items
-
Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model
by: Liu, Haogeng, et al.
Published: (2024) -
Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models
by: Liu, Zijian, et al.
Published: (2026) -
Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
by: Zhou, Haoran, et al.
Published: (2025) -
Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation
by: Wu, Yebo, et al.
Published: (2026) -
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
by: Li, Bin, et al.
Published: (2025)