Visual Hallucinations of Multi-modal Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Wen, Liu, Hongbin, Guo, Minxin, Gong, Neil Zhenqiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Automatically Generating Visual Hallucination Test Cases for Multimodal Large Language Models
di: Liu, Zhongye, et al.
Pubblicazione: (2024)
di: Liu, Zhongye, et al.
Pubblicazione: (2024)
Refusing Safe Prompts for Multi-modal Large Language Models
di: Shao, Zedian, et al.
Pubblicazione: (2024)
di: Shao, Zedian, et al.
Pubblicazione: (2024)
Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection
di: Shao, Zedian, et al.
Pubblicazione: (2026)
di: Shao, Zedian, et al.
Pubblicazione: (2026)
VIAssist: Adapting Multi-modal Large Language Models for Users with Visual Impairments
di: Yang, Bufang, et al.
Pubblicazione: (2024)
di: Yang, Bufang, et al.
Pubblicazione: (2024)
Securing Visually-Aware Recommender Systems: An Adversarial Image Reconstruction and Detection Framework
di: Yin, Minglei, et al.
Pubblicazione: (2023)
di: Yin, Minglei, et al.
Pubblicazione: (2023)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
di: Park, Yeji, et al.
Pubblicazione: (2024)
di: Park, Yeji, et al.
Pubblicazione: (2024)
Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
di: Tang, Kai, et al.
Pubblicazione: (2025)
di: Tang, Kai, et al.
Pubblicazione: (2025)
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
EditTrack: Detecting and Attributing AI-assisted Image Editing
di: Jiang, Zhengyuan, et al.
Pubblicazione: (2025)
di: Jiang, Zhengyuan, et al.
Pubblicazione: (2025)
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
di: Lee, Jihoon, et al.
Pubblicazione: (2025)
di: Lee, Jihoon, et al.
Pubblicazione: (2025)
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
di: He, Hulingxiao, et al.
Pubblicazione: (2025)
di: He, Hulingxiao, et al.
Pubblicazione: (2025)
Mudjacking: Patching Backdoor Vulnerabilities in Foundation Models
di: Liu, Hongbin, et al.
Pubblicazione: (2024)
di: Liu, Hongbin, et al.
Pubblicazione: (2024)
Causal Decoding for Hallucination-Resistant Multimodal Large Language Models
di: Tan, Shiwei, et al.
Pubblicazione: (2026)
di: Tan, Shiwei, et al.
Pubblicazione: (2026)
Generative Multi-modal Models are Good Class-Incremental Learners
di: Cao, Xusheng, et al.
Pubblicazione: (2024)
di: Cao, Xusheng, et al.
Pubblicazione: (2024)
Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models
di: Li, Shengzhi, et al.
Pubblicazione: (2024)
di: Li, Shengzhi, et al.
Pubblicazione: (2024)
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
di: Liu, Dongyang, et al.
Pubblicazione: (2024)
di: Liu, Dongyang, et al.
Pubblicazione: (2024)
Watermark-based Attribution of AI-Generated Content
di: Jiang, Zhengyuan, et al.
Pubblicazione: (2024)
di: Jiang, Zhengyuan, et al.
Pubblicazione: (2024)
Woodpecker: Hallucination Correction for Multimodal Large Language Models
di: Yin, Shukang, et al.
Pubblicazione: (2023)
di: Yin, Shukang, et al.
Pubblicazione: (2023)
Evaluating Hallucination in Large Vision-Language Models based on Context-Aware Object Similarities
di: Datta, Shounak, et al.
Pubblicazione: (2025)
di: Datta, Shounak, et al.
Pubblicazione: (2025)
Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental Learning
di: Cao, Xusheng, et al.
Pubblicazione: (2025)
di: Cao, Xusheng, et al.
Pubblicazione: (2025)
Efficient Contrastive Decoding with Probabilistic Hallucination Detection - Mitigating Hallucinations in Large Vision Language Models -
di: Fieback, Laura, et al.
Pubblicazione: (2025)
di: Fieback, Laura, et al.
Pubblicazione: (2025)
WebInject: Prompt Injection Attack to Web Agents
di: Wang, Xilong, et al.
Pubblicazione: (2025)
di: Wang, Xilong, et al.
Pubblicazione: (2025)
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
di: Wu, Junfei, et al.
Pubblicazione: (2024)
di: Wu, Junfei, et al.
Pubblicazione: (2024)
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
di: Kaul, Prannay, et al.
Pubblicazione: (2024)
di: Kaul, Prannay, et al.
Pubblicazione: (2024)
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
di: Sun, Hai-Long, et al.
Pubblicazione: (2025)
di: Sun, Hai-Long, et al.
Pubblicazione: (2025)
VideoMarkBench: Benchmarking Robustness of Video Watermarking
di: Jiang, Zhengyuan, et al.
Pubblicazione: (2025)
di: Jiang, Zhengyuan, et al.
Pubblicazione: (2025)
HalluRNN: Mitigating Hallucinations via Recurrent Cross-Layer Reasoning in Large Vision-Language Models
di: Yu, Le, et al.
Pubblicazione: (2025)
di: Yu, Le, et al.
Pubblicazione: (2025)
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
di: Wu, Junjie, et al.
Pubblicazione: (2024)
di: Wu, Junjie, et al.
Pubblicazione: (2024)
Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
di: Hong, Yunqi, et al.
Pubblicazione: (2025)
di: Hong, Yunqi, et al.
Pubblicazione: (2025)
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
di: Wang, Yubo, et al.
Pubblicazione: (2024)
di: Wang, Yubo, et al.
Pubblicazione: (2024)
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
di: Xiao, Wenyi, et al.
Pubblicazione: (2024)
di: Xiao, Wenyi, et al.
Pubblicazione: (2024)
CorruptEncoder: Data Poisoning based Backdoor Attacks to Contrastive Learning
di: Zhang, Jinghuai, et al.
Pubblicazione: (2022)
di: Zhang, Jinghuai, et al.
Pubblicazione: (2022)
Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models
di: Dong, Yuhao, et al.
Pubblicazione: (2026)
di: Dong, Yuhao, et al.
Pubblicazione: (2026)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
di: Huang, Po-Hsuan, et al.
Pubblicazione: (2024)
di: Huang, Po-Hsuan, et al.
Pubblicazione: (2024)
Exploring the Transferability of Visual Prompting for Multimodal Large Language Models
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
Prompting Medical Large Vision-Language Models to Diagnose Pathologies by Visual Question Answering
di: Guo, Danfeng, et al.
Pubblicazione: (2024)
di: Guo, Danfeng, et al.
Pubblicazione: (2024)
Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data
di: Yoon, Heegeon, et al.
Pubblicazione: (2026)
di: Yoon, Heegeon, et al.
Pubblicazione: (2026)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
di: Mena, Francisco, et al.
Pubblicazione: (2025)
di: Mena, Francisco, et al.
Pubblicazione: (2025)
Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance
di: Zhao, Linxi, et al.
Pubblicazione: (2024)
di: Zhao, Linxi, et al.
Pubblicazione: (2024)
Skip \n: A Simple Method to Reduce Hallucination in Large Vision-Language Models
di: Han, Zongbo, et al.
Pubblicazione: (2024)
di: Han, Zongbo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Automatically Generating Visual Hallucination Test Cases for Multimodal Large Language Models
di: Liu, Zhongye, et al.
Pubblicazione: (2024) -
Refusing Safe Prompts for Multi-modal Large Language Models
di: Shao, Zedian, et al.
Pubblicazione: (2024) -
Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection
di: Shao, Zedian, et al.
Pubblicazione: (2026) -
VIAssist: Adapting Multi-modal Large Language Models for Users with Visual Impairments
di: Yang, Bufang, et al.
Pubblicazione: (2024) -
Securing Visually-Aware Recommender Systems: An Adversarial Image Reconstruction and Detection Framework
di: Yin, Minglei, et al.
Pubblicazione: (2023)