Multi-Object Hallucination in Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Xuweiyi, Ma, Ziqiao, Zhang, Xuejun, Xu, Sihan, Qian, Shengyi, Yang, Jianing, Fouhey, David F., Chai, Joyce |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination
por: Yang, Jianing, et al.
Publicado: (2024)
por: Yang, Jianing, et al.
Publicado: (2024)
World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models
por: Ma, Ziqiao, et al.
Publicado: (2023)
por: Ma, Ziqiao, et al.
Publicado: (2023)
Next-Embedding Prediction Makes Strong Vision Learners
por: Xu, Sihan, et al.
Publicado: (2025)
por: Xu, Sihan, et al.
Publicado: (2025)
SAB3R: Semantic-Augmented Backbone in 3D Reconstruction
por: Chen, Xuweiyi, et al.
Publicado: (2025)
por: Chen, Xuweiyi, et al.
Publicado: (2025)
Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
por: Zhang, Zheyuan, et al.
Publicado: (2024)
por: Zhang, Zheyuan, et al.
Publicado: (2024)
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
por: Zhang, Yue, et al.
Publicado: (2024)
por: Zhang, Yue, et al.
Publicado: (2024)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
por: Chen, Boqi, et al.
Publicado: (2026)
por: Chen, Boqi, et al.
Publicado: (2026)
The Mechanistic Emergence of Symbol Grounding in Language Models
por: Wu, Shuyu, et al.
Publicado: (2025)
por: Wu, Shuyu, et al.
Publicado: (2025)
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
por: Zhang, Yichi, et al.
Publicado: (2024)
por: Zhang, Yichi, et al.
Publicado: (2024)
4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
por: Ma, Ziqiao, et al.
Publicado: (2025)
por: Ma, Ziqiao, et al.
Publicado: (2025)
Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation
por: Ma, Ziqiao, et al.
Publicado: (2025)
por: Ma, Ziqiao, et al.
Publicado: (2025)
Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
por: Ma, Ruiqi, et al.
Publicado: (2025)
por: Ma, Ruiqi, et al.
Publicado: (2025)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
por: Shang, Yuying, et al.
Publicado: (2024)
por: Shang, Yuying, et al.
Publicado: (2024)
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
por: Lovenia, Holy, et al.
Publicado: (2023)
por: Lovenia, Holy, et al.
Publicado: (2023)
DriVLMe: Enhancing LLM-based Autonomous Driving Agents with Embodied and Social Experiences
por: Huang, Yidong, et al.
Publicado: (2024)
por: Huang, Yidong, et al.
Publicado: (2024)
UniCtrl: Improving the Spatiotemporal Consistency of Text-to-Video Diffusion Models via Training-Free Unified Attention Control
por: Xia, Tian, et al.
Publicado: (2024)
por: Xia, Tian, et al.
Publicado: (2024)
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
por: Geigle, Gregor, et al.
Publicado: (2024)
por: Geigle, Gregor, et al.
Publicado: (2024)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
por: Jing, Liqiang, et al.
Publicado: (2025)
por: Jing, Liqiang, et al.
Publicado: (2025)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
por: Zhou, Yiyang, et al.
Publicado: (2023)
por: Zhou, Yiyang, et al.
Publicado: (2023)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
por: An, Wenbin, et al.
Publicado: (2024)
por: An, Wenbin, et al.
Publicado: (2024)
ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
por: Chen, Junzhe, et al.
Publicado: (2024)
por: Chen, Junzhe, et al.
Publicado: (2024)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
por: Lu, Yifan, et al.
Publicado: (2025)
por: Lu, Yifan, et al.
Publicado: (2025)
CycleNet: Rethinking Cycle Consistency in Text-Guided Diffusion for Image Manipulation
por: Xu, Sihan, et al.
Publicado: (2023)
por: Xu, Sihan, et al.
Publicado: (2023)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
por: Seo, Hoigi, et al.
Publicado: (2025)
por: Seo, Hoigi, et al.
Publicado: (2025)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
por: Oh, Hongseok, et al.
Publicado: (2025)
por: Oh, Hongseok, et al.
Publicado: (2025)
Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use
por: Xi, Jiajun, et al.
Publicado: (2024)
por: Xi, Jiajun, et al.
Publicado: (2024)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
por: Wang, Zihu, et al.
Publicado: (2025)
por: Wang, Zihu, et al.
Publicado: (2025)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
por: Li, Bin, et al.
Publicado: (2025)
por: Li, Bin, et al.
Publicado: (2025)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
por: Yu, Shoubin, et al.
Publicado: (2025)
por: Yu, Shoubin, et al.
Publicado: (2025)
Benchmarking Deflection and Hallucination in Large Vision-Language Models
por: Moratelli, Nicholas, et al.
Publicado: (2026)
por: Moratelli, Nicholas, et al.
Publicado: (2026)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
por: Zhong, Weihong, et al.
Publicado: (2024)
por: Zhong, Weihong, et al.
Publicado: (2024)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
por: Wang, Chao, et al.
Publicado: (2025)
por: Wang, Chao, et al.
Publicado: (2025)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
por: Ye, Zekai, et al.
Publicado: (2025)
por: Ye, Zekai, et al.
Publicado: (2025)
ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
por: Wan, Zifu, et al.
Publicado: (2025)
por: Wan, Zifu, et al.
Publicado: (2025)
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
por: Zhang, Ce, et al.
Publicado: (2025)
por: Zhang, Ce, et al.
Publicado: (2025)
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
por: Wang, Weihang, et al.
Publicado: (2025)
por: Wang, Weihang, et al.
Publicado: (2025)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
por: Woo, Sangmin, et al.
Publicado: (2025)
por: Woo, Sangmin, et al.
Publicado: (2025)
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties
por: Yu, Keunwoo Peter, et al.
Publicado: (2023)
por: Yu, Keunwoo Peter, et al.
Publicado: (2023)
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
por: Wu, Xiyang, et al.
Publicado: (2024)
por: Wu, Xiyang, et al.
Publicado: (2024)
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
por: Chang, Yue, et al.
Publicado: (2024)
por: Chang, Yue, et al.
Publicado: (2024)
Ejemplares similares
-
3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination
por: Yang, Jianing, et al.
Publicado: (2024) -
World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models
por: Ma, Ziqiao, et al.
Publicado: (2023) -
Next-Embedding Prediction Makes Strong Vision Learners
por: Xu, Sihan, et al.
Publicado: (2025) -
SAB3R: Semantic-Augmented Backbone in 3D Reconstruction
por: Chen, Xuweiyi, et al.
Publicado: (2025) -
Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
por: Zhang, Zheyuan, et al.
Publicado: (2024)