IKOD: Mitigating Visual Attention Degradation in Large Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Jiabing, Cui, Chenhang, Zhou, Yiyang, Chen, Yixiang, Xia, Peng, Wei, Ying, Yu, Tao, Huang, Yan, Wang, Liang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
di: Zhou, Yiyang, et al.
Pubblicazione: (2023)
di: Zhou, Yiyang, et al.
Pubblicazione: (2023)
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
di: Zhou, Yiyang, et al.
Pubblicazione: (2024)
di: Zhou, Yiyang, et al.
Pubblicazione: (2024)
Improving Alignment in LVLMs with Debiased Self-Judgment
di: Yang, Sihan, et al.
Pubblicazione: (2025)
di: Yang, Sihan, et al.
Pubblicazione: (2025)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
di: Xia, Peng, et al.
Pubblicazione: (2024)
di: Xia, Peng, et al.
Pubblicazione: (2024)
Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
di: Cui, Chenhang, et al.
Pubblicazione: (2024)
di: Cui, Chenhang, et al.
Pubblicazione: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
di: Li, Bin, et al.
Pubblicazione: (2025)
di: Li, Bin, et al.
Pubblicazione: (2025)
LaPA$^2$: Length-Aware Prefix and Prompt Attention Augmentation for Long-Form Controllable Text Generation
di: Yang, Jiabing, et al.
Pubblicazione: (2025)
di: Yang, Jiabing, et al.
Pubblicazione: (2025)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
di: Wang, Zihu, et al.
Pubblicazione: (2025)
di: Wang, Zihu, et al.
Pubblicazione: (2025)
Calibrated Self-Rewarding Vision Language Models
di: Zhou, Yiyang, et al.
Pubblicazione: (2024)
di: Zhou, Yiyang, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
di: Zhu, Xingyu, et al.
Pubblicazione: (2026)
di: Zhu, Xingyu, et al.
Pubblicazione: (2026)
UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models
di: Yang, Jiabing, et al.
Pubblicazione: (2026)
di: Yang, Jiabing, et al.
Pubblicazione: (2026)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
di: Zhu, Younan, et al.
Pubblicazione: (2025)
di: Zhu, Younan, et al.
Pubblicazione: (2025)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
di: Li, Qiming, et al.
Pubblicazione: (2026)
di: Li, Qiming, et al.
Pubblicazione: (2026)
SDCD: Structure-Disrupted Contrastive Decoding for Mitigating Hallucinations in Large Vision-Language Models
di: Xia, Yuxuan, et al.
Pubblicazione: (2026)
di: Xia, Yuxuan, et al.
Pubblicazione: (2026)
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025)
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025)
Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models
di: Shi, Enyi, et al.
Pubblicazione: (2026)
di: Shi, Enyi, et al.
Pubblicazione: (2026)
Mitigating Multilingual Hallucination in Large Vision-Language Models
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
di: An, Wenbin, et al.
Pubblicazione: (2024)
di: An, Wenbin, et al.
Pubblicazione: (2024)
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
di: Li, Chenxi, et al.
Pubblicazione: (2025)
di: Li, Chenxi, et al.
Pubblicazione: (2025)
Rethinking Causal Mask Attention for Vision-Language Inference
di: Pei, Xiaohuan, et al.
Pubblicazione: (2025)
di: Pei, Xiaohuan, et al.
Pubblicazione: (2025)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
di: Lee, Yi-Lun, et al.
Pubblicazione: (2024)
di: Lee, Yi-Lun, et al.
Pubblicazione: (2024)
MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
di: Yang, Chengxu, et al.
Pubblicazione: (2026)
di: Yang, Chengxu, et al.
Pubblicazione: (2026)
VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
di: Yang, Sihan, et al.
Pubblicazione: (2025)
di: Yang, Sihan, et al.
Pubblicazione: (2025)
Large Vision-Language Models Get Lost in Attention
di: Xi, Gongli, et al.
Pubblicazione: (2026)
di: Xi, Gongli, et al.
Pubblicazione: (2026)
RCP: Representation Consistency Pruner for Mitigating Distribution Shift in Large Vision-Language Models
di: Zhang, Jianwei, et al.
Pubblicazione: (2026)
di: Zhang, Jianwei, et al.
Pubblicazione: (2026)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
di: Woo, Sangmin, et al.
Pubblicazione: (2025)
di: Woo, Sangmin, et al.
Pubblicazione: (2025)
Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework
di: Jia, Hongrui, et al.
Pubblicazione: (2025)
di: Jia, Hongrui, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
Attention Prompting on Image for Large Vision-Language Models
di: Yu, Runpeng, et al.
Pubblicazione: (2024)
di: Yu, Runpeng, et al.
Pubblicazione: (2024)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
di: Ye, Zekai, et al.
Pubblicazione: (2025)
di: Ye, Zekai, et al.
Pubblicazione: (2025)
LAPT: Label-driven Automated Prompt Tuning for OOD Detection with Vision-Language Models
di: Zhang, Yabin, et al.
Pubblicazione: (2024)
di: Zhang, Yabin, et al.
Pubblicazione: (2024)
RobustVisRAG: Causality-Aware Vision-Based Retrieval-Augmented Generation under Visual Degradations
di: Chen, I-Hsiang, et al.
Pubblicazione: (2026)
di: Chen, I-Hsiang, et al.
Pubblicazione: (2026)
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
di: Arif, Kazi Hasan Ibn, et al.
Pubblicazione: (2025)
di: Arif, Kazi Hasan Ibn, et al.
Pubblicazione: (2025)
Adversarial Robustness for Visual Grounding of Multimodal Large Language Models
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
di: Li, Qiming, et al.
Pubblicazione: (2025)
di: Li, Qiming, et al.
Pubblicazione: (2025)
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models
di: Hong, Sujung, et al.
Pubblicazione: (2026)
di: Hong, Sujung, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
di: Zhou, Yiyang, et al.
Pubblicazione: (2023) -
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
di: Zhou, Yiyang, et al.
Pubblicazione: (2024) -
Improving Alignment in LVLMs with Debiased Self-Judgment
di: Yang, Sihan, et al.
Pubblicazione: (2025) -
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
di: Xia, Peng, et al.
Pubblicazione: (2024) -
Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
di: Cui, Chenhang, et al.
Pubblicazione: (2024)