CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Qiming, Ye, Zekai, Feng, Xiaocheng, Zhong, Weihong, Qin, Libo, Chen, Ruihan, Li, Baohang, Jiang, Kui, Wang, Yaowei, Liu, Ting, Qin, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
by: Li, Qiming, et al.
Published: (2026)
by: Li, Qiming, et al.
Published: (2026)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
by: Ye, Zekai, et al.
Published: (2025)
by: Ye, Zekai, et al.
Published: (2025)
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
by: Li, Qiming, et al.
Published: (2025)
by: Li, Qiming, et al.
Published: (2025)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
by: Zhong, Weihong, et al.
Published: (2024)
by: Zhong, Weihong, et al.
Published: (2024)
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
by: Li, Qiming, et al.
Published: (2025)
by: Li, Qiming, et al.
Published: (2025)
Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models
by: Ye, Zekai, et al.
Published: (2026)
by: Ye, Zekai, et al.
Published: (2026)
Aligning Translation-Specific Understanding to General Understanding in Large Language Models
by: Huang, Yichong, et al.
Published: (2024)
by: Huang, Yichong, et al.
Published: (2024)
Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel Collaboration
by: Huang, Yichong, et al.
Published: (2024)
by: Huang, Yichong, et al.
Published: (2024)
Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification
by: Sun, Han, et al.
Published: (2026)
by: Sun, Han, et al.
Published: (2026)
Unveiling Entity-Level Unlearning for Large Language Models: A Comprehensive Analysis
by: Ma, Weitao, et al.
Published: (2024)
by: Ma, Weitao, et al.
Published: (2024)
Culture-Aware Machine Translation in Large Language Models: Benchmarking and Investigation
by: Yuan, Zekun, et al.
Published: (2026)
by: Yuan, Zekun, et al.
Published: (2026)
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
by: Huang, Lei, et al.
Published: (2023)
by: Huang, Lei, et al.
Published: (2023)
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
by: Chen, Ruihan, et al.
Published: (2025)
by: Chen, Ruihan, et al.
Published: (2025)
Mitigating Image Captioning Hallucinations in Vision-Language Models
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs
by: Yu, Liu, et al.
Published: (2025)
by: Yu, Liu, et al.
Published: (2025)
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play
by: Feng, Xiachong, et al.
Published: (2026)
by: Feng, Xiachong, et al.
Published: (2026)
Relay Decoding: Concatenating Large Language Models for Machine Translation
by: Fu, Chengpeng, et al.
Published: (2024)
by: Fu, Chengpeng, et al.
Published: (2024)
Can Large Language Models Simulate Human Cognition Beyond Behavioral Imitation?
by: Gu, Yuxuan, et al.
Published: (2026)
by: Gu, Yuxuan, et al.
Published: (2026)
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding
by: Zhao, Liang, et al.
Published: (2023)
by: Zhao, Liang, et al.
Published: (2023)
GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News Summarization
by: Ye, Yangfan, et al.
Published: (2024)
by: Ye, Yangfan, et al.
Published: (2024)
Mitigating Object Hallucination via Concentric Causal Attention
by: Xing, Yun, et al.
Published: (2024)
by: Xing, Yun, et al.
Published: (2024)
Mitigating Object Hallucinations via Sentence-Level Early Intervention
by: Peng, Shangpin, et al.
Published: (2025)
by: Peng, Shangpin, et al.
Published: (2025)
One for All: Update Parameterized Knowledge Across Multiple Models
by: Ma, Weitao, et al.
Published: (2025)
by: Ma, Weitao, et al.
Published: (2025)
PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
by: Feng, Xiachong, et al.
Published: (2026)
by: Feng, Xiachong, et al.
Published: (2026)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
Extending Context Window of Large Language Models from a Distributional Perspective
by: Wu, Yingsheng, et al.
Published: (2024)
by: Wu, Yingsheng, et al.
Published: (2024)
Discrete Modeling via Boundary Conditional Diffusion Processes
by: Gu, Yuxuan, et al.
Published: (2024)
by: Gu, Yuxuan, et al.
Published: (2024)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
by: Zhu, Younan, et al.
Published: (2025)
by: Zhu, Younan, et al.
Published: (2025)
HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
by: Yuan, Fan, et al.
Published: (2024)
by: Yuan, Fan, et al.
Published: (2024)
FroM: Frobenius Norm-Based Data-Free Adaptive Model Merging
by: Li, Zijian, et al.
Published: (2025)
by: Li, Zijian, et al.
Published: (2025)
Mitigating Open-Vocabulary Caption Hallucinations
by: Ben-Kish, Assaf, et al.
Published: (2023)
by: Ben-Kish, Assaf, et al.
Published: (2023)
Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
by: Zhang, Chengsheng, et al.
Published: (2026)
by: Zhang, Chengsheng, et al.
Published: (2026)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
by: Cheng, Zihui, et al.
Published: (2024)
by: Cheng, Zihui, et al.
Published: (2024)
Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective
by: Yu, Xinmiao, et al.
Published: (2024)
by: Yu, Xinmiao, et al.
Published: (2024)
Look Closer! An Adversarial Parametric Editing Framework for Hallucination Mitigation in VLMs
by: Hu, Jiayu, et al.
Published: (2025)
by: Hu, Jiayu, et al.
Published: (2025)
Length Controlled Generation for Black-box LLMs
by: Gu, Yuxuan, et al.
Published: (2024)
by: Gu, Yuxuan, et al.
Published: (2024)
ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
by: Chen, Junzhe, et al.
Published: (2024)
by: Chen, Junzhe, et al.
Published: (2024)
Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration
by: Xu, Yuanzhi, et al.
Published: (2026)
by: Xu, Yuanzhi, et al.
Published: (2026)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
by: Lu, Yifan, et al.
Published: (2025)
by: Lu, Yifan, et al.
Published: (2025)
Similar Items
-
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
by: Li, Qiming, et al.
Published: (2026) -
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
by: Ye, Zekai, et al.
Published: (2025) -
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
by: Li, Qiming, et al.
Published: (2025) -
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
by: Zhong, Weihong, et al.
Published: (2024) -
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
by: Li, Qiming, et al.
Published: (2025)