Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Yiyang, Cui, Chenhang, Yoon, Jaehong, Zhang, Linjun, Deng, Zhun, Finn, Chelsea, Bansal, Mohit, Yao, Huaxiu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
di: Zhou, Yiyang, et al.
Pubblicazione: (2024)
di: Zhou, Yiyang, et al.
Pubblicazione: (2024)
Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
di: Cui, Chenhang, et al.
Pubblicazione: (2024)
di: Cui, Chenhang, et al.
Pubblicazione: (2024)
Calibrated Self-Rewarding Vision Language Models
di: Zhou, Yiyang, et al.
Pubblicazione: (2024)
di: Zhou, Yiyang, et al.
Pubblicazione: (2024)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
di: Sung, Yi-Lin, et al.
Pubblicazione: (2023)
di: Sung, Yi-Lin, et al.
Pubblicazione: (2023)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
di: Yu, Shoubin, et al.
Pubblicazione: (2024)
di: Yu, Shoubin, et al.
Pubblicazione: (2024)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
Multimodal Representation Learning by Alternating Unimodal Adaptation
di: Zhang, Xiaohui, et al.
Pubblicazione: (2023)
di: Zhang, Xiaohui, et al.
Pubblicazione: (2023)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
di: Yu, Shoubin, et al.
Pubblicazione: (2026)
di: Yu, Shoubin, et al.
Pubblicazione: (2026)
Improving Alignment in LVLMs with Debiased Self-Judgment
di: Yang, Sihan, et al.
Pubblicazione: (2025)
di: Yang, Sihan, et al.
Pubblicazione: (2025)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
di: Xia, Peng, et al.
Pubblicazione: (2024)
di: Xia, Peng, et al.
Pubblicazione: (2024)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
di: Lee, Daeun, et al.
Pubblicazione: (2024)
di: Lee, Daeun, et al.
Pubblicazione: (2024)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
di: Lee, Daeun, et al.
Pubblicazione: (2025)
di: Lee, Daeun, et al.
Pubblicazione: (2025)
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
di: Liang, Yiming, et al.
Pubblicazione: (2026)
di: Liang, Yiming, et al.
Pubblicazione: (2026)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
Hierarchy-Aware Multimodal Unlearning for Medical AI
di: Wu, Fengli, et al.
Pubblicazione: (2025)
di: Wu, Fengli, et al.
Pubblicazione: (2025)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
di: Wang, Zun, et al.
Pubblicazione: (2024)
di: Wang, Zun, et al.
Pubblicazione: (2024)
Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance
di: Zhao, Linxi, et al.
Pubblicazione: (2024)
di: Zhao, Linxi, et al.
Pubblicazione: (2024)
Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
di: Ma, Ruiqi, et al.
Pubblicazione: (2025)
di: Ma, Ruiqi, et al.
Pubblicazione: (2025)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
di: Woo, Sangmin, et al.
Pubblicazione: (2025)
di: Woo, Sangmin, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
di: Lu, Yifan, et al.
Pubblicazione: (2025)
di: Lu, Yifan, et al.
Pubblicazione: (2025)
ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
di: Chen, Junzhe, et al.
Pubblicazione: (2024)
di: Chen, Junzhe, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
di: An, Wenbin, et al.
Pubblicazione: (2024)
di: An, Wenbin, et al.
Pubblicazione: (2024)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
di: Li, Jialu, et al.
Pubblicazione: (2025)
di: Li, Jialu, et al.
Pubblicazione: (2025)
Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
di: Zheng, Haohan, et al.
Pubblicazione: (2025)
di: Zheng, Haohan, et al.
Pubblicazione: (2025)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
di: Li, Jialu, et al.
Pubblicazione: (2024)
di: Li, Jialu, et al.
Pubblicazione: (2024)
Mitigating Multilingual Hallucination in Large Vision-Language Models
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
di: Chang, Yue, et al.
Pubblicazione: (2024)
di: Chang, Yue, et al.
Pubblicazione: (2024)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
di: Pothiraj, Atin, et al.
Pubblicazione: (2025)
di: Pothiraj, Atin, et al.
Pubblicazione: (2025)
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
di: Wang, Weihang, et al.
Pubblicazione: (2025)
di: Wang, Weihang, et al.
Pubblicazione: (2025)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
di: Ren, Lingfeng, et al.
Pubblicazione: (2026)
di: Ren, Lingfeng, et al.
Pubblicazione: (2026)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
di: Ye, Zekai, et al.
Pubblicazione: (2025)
di: Ye, Zekai, et al.
Pubblicazione: (2025)
First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
di: Ha, Jiwoo, et al.
Pubblicazione: (2026)
di: Ha, Jiwoo, et al.
Pubblicazione: (2026)
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
di: Agrawal, Aakriti, et al.
Pubblicazione: (2025)
di: Agrawal, Aakriti, et al.
Pubblicazione: (2025)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
di: Li, Bin, et al.
Pubblicazione: (2025)
di: Li, Bin, et al.
Pubblicazione: (2025)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
di: Zhou, Yiyang, et al.
Pubblicazione: (2024) -
Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
di: Cui, Chenhang, et al.
Pubblicazione: (2024) -
Calibrated Self-Rewarding Vision Language Models
di: Zhou, Yiyang, et al.
Pubblicazione: (2024) -
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
di: Sung, Yi-Lin, et al.
Pubblicazione: (2023) -
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
di: Yu, Shoubin, et al.
Pubblicazione: (2024)