CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Qiming, Ye, Zekai, Feng, Xiaocheng, Zhong, Weihong, Qin, Libo, Chen, Ruihan, Huang, Lei, Li, Baohang, Jiang, Kui, Wang, Yaowei, Liu, Ting, Qin, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
by: Li, Qiming, et al.
Published: (2025)
by: Li, Qiming, et al.
Published: (2025)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
by: Ye, Zekai, et al.
Published: (2025)
by: Ye, Zekai, et al.
Published: (2025)
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
by: Li, Qiming, et al.
Published: (2025)
by: Li, Qiming, et al.
Published: (2025)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
by: Zhong, Weihong, et al.
Published: (2024)
by: Zhong, Weihong, et al.
Published: (2024)
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
by: Li, Qiming, et al.
Published: (2025)
by: Li, Qiming, et al.
Published: (2025)
Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models
by: Ye, Zekai, et al.
Published: (2026)
by: Ye, Zekai, et al.
Published: (2026)
Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification
by: Sun, Han, et al.
Published: (2026)
by: Sun, Han, et al.
Published: (2026)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
by: Wang, Zihu, et al.
Published: (2025)
by: Wang, Zihu, et al.
Published: (2025)
Aligning Translation-Specific Understanding to General Understanding in Large Language Models
by: Huang, Yichong, et al.
Published: (2024)
by: Huang, Yichong, et al.
Published: (2024)
Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel Collaboration
by: Huang, Yichong, et al.
Published: (2024)
by: Huang, Yichong, et al.
Published: (2024)
Mitigating Image Captioning Hallucinations in Vision-Language Models
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
by: Yin, Jianghao, et al.
Published: (2026)
by: Yin, Jianghao, et al.
Published: (2026)
Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations
by: Chen, Boxu, et al.
Published: (2025)
by: Chen, Boxu, et al.
Published: (2025)
Culture-Aware Machine Translation in Large Language Models: Benchmarking and Investigation
by: Yuan, Zekun, et al.
Published: (2026)
by: Yuan, Zekun, et al.
Published: (2026)
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
by: Chen, Ruihan, et al.
Published: (2025)
by: Chen, Ruihan, et al.
Published: (2025)
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
by: Huang, Lei, et al.
Published: (2023)
by: Huang, Lei, et al.
Published: (2023)
Unveiling Entity-Level Unlearning for Large Language Models: A Comprehensive Analysis
by: Ma, Weitao, et al.
Published: (2024)
by: Ma, Weitao, et al.
Published: (2024)
Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
by: Wu, Jialin, et al.
Published: (2026)
by: Wu, Jialin, et al.
Published: (2026)
Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement
by: Qin, Zhenxin, et al.
Published: (2026)
by: Qin, Zhenxin, et al.
Published: (2026)
Mitigating Object Hallucination via Concentric Causal Attention
by: Xing, Yun, et al.
Published: (2024)
by: Xing, Yun, et al.
Published: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective
by: Yu, Xinmiao, et al.
Published: (2024)
by: Yu, Xinmiao, et al.
Published: (2024)
Relay Decoding: Concatenating Large Language Models for Machine Translation
by: Fu, Chengpeng, et al.
Published: (2024)
by: Fu, Chengpeng, et al.
Published: (2024)
MHSA: A Lightweight Framework for Mitigating Hallucinations via Steered Attention in LVLMs
by: Ding, Wei, et al.
Published: (2026)
by: Ding, Wei, et al.
Published: (2026)
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play
by: Feng, Xiachong, et al.
Published: (2026)
by: Feng, Xiachong, et al.
Published: (2026)
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
by: Xu, Le, et al.
Published: (2025)
by: Xu, Le, et al.
Published: (2025)
Can Large Language Models Simulate Human Cognition Beyond Behavioral Imitation?
by: Gu, Yuxuan, et al.
Published: (2026)
by: Gu, Yuxuan, et al.
Published: (2026)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
by: Zhu, Younan, et al.
Published: (2025)
by: Zhu, Younan, et al.
Published: (2025)
HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
by: Yuan, Fan, et al.
Published: (2024)
by: Yuan, Fan, et al.
Published: (2024)
COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
by: Sahay, Kenji, et al.
Published: (2025)
by: Sahay, Kenji, et al.
Published: (2025)
Mitigating Open-Vocabulary Caption Hallucinations
by: Ben-Kish, Assaf, et al.
Published: (2023)
by: Ben-Kish, Assaf, et al.
Published: (2023)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
by: Zhang, Yuanhong, et al.
Published: (2026)
by: Zhang, Yuanhong, et al.
Published: (2026)
Look Closer! An Adversarial Parametric Editing Framework for Hallucination Mitigation in VLMs
by: Hu, Jiayu, et al.
Published: (2025)
by: Hu, Jiayu, et al.
Published: (2025)
Energy-Guided Decoding for Object Hallucination Mitigation
by: Liu, Xixi, et al.
Published: (2025)
by: Liu, Xixi, et al.
Published: (2025)
Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration
by: Xu, Yuanzhi, et al.
Published: (2026)
by: Xu, Yuanzhi, et al.
Published: (2026)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding
by: Zhao, Liang, et al.
Published: (2023)
by: Zhao, Liang, et al.
Published: (2023)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
by: Lu, Yifan, et al.
Published: (2025)
by: Lu, Yifan, et al.
Published: (2025)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
by: Zou, Zhengtao, et al.
Published: (2025)
by: Zou, Zhengtao, et al.
Published: (2025)
Similar Items
-
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
by: Li, Qiming, et al.
Published: (2025) -
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
by: Ye, Zekai, et al.
Published: (2025) -
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
by: Li, Qiming, et al.
Published: (2025) -
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
by: Zhong, Weihong, et al.
Published: (2024) -
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
by: Li, Qiming, et al.
Published: (2025)