A Survey of Multimodal Hallucination Evaluation and Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zhiyuan, Min, Yuecong, Zhang, Jie, Yan, Bei, Wang, Jiahao, Wang, Xiaozhen, Shan, Shiguang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
by: Yan, Bei, et al.
Published: (2025)
by: Yan, Bei, et al.
Published: (2025)
ACT Now: Preempting LVLM Hallucinations via Adaptive Context Integration
by: Yan, Bei, et al.
Published: (2026)
by: Yan, Bei, et al.
Published: (2026)
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
by: Yang, Junqi, et al.
Published: (2026)
by: Yang, Junqi, et al.
Published: (2026)
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
by: Yan, Bei, et al.
Published: (2024)
by: Yan, Bei, et al.
Published: (2024)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
by: Yan, Bei, et al.
Published: (2024)
by: Yan, Bei, et al.
Published: (2024)
T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models
by: Li, Changzhen, et al.
Published: (2025)
by: Li, Changzhen, et al.
Published: (2025)
Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Rethinking the Evaluation of Out-of-Distribution Detection: A Sorites Paradox
by: Long, Xingming, et al.
Published: (2024)
by: Long, Xingming, et al.
Published: (2024)
VOPE: Revisiting Hallucination of Vision-Language Models in Voluntary Imagination Task
by: Long, Xingming, et al.
Published: (2025)
by: Long, Xingming, et al.
Published: (2025)
Dynamic Attention Analysis for Backdoor Detection in Text-to-Image Diffusion Models
by: Wang, Zhongqi, et al.
Published: (2025)
by: Wang, Zhongqi, et al.
Published: (2025)
Assimilation Matters: Model-level Backdoor Detection in Vision-Language Pretrained Models
by: Wang, Zhongqi, et al.
Published: (2025)
by: Wang, Zhongqi, et al.
Published: (2025)
REVAL: A Comprehension Evaluation on Reliability and Values of Large Vision-Language Models
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Semantic or Covariate? A Study on the Intractable Case of Out-of-Distribution Detection
by: Long, Xingming, et al.
Published: (2024)
by: Long, Xingming, et al.
Published: (2024)
Generalized Face Liveness Detection via De-fake Face Generator
by: Long, Xingming, et al.
Published: (2024)
by: Long, Xingming, et al.
Published: (2024)
EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy
by: Ge, Xuanyu, et al.
Published: (2026)
by: Ge, Xuanyu, et al.
Published: (2026)
T2IShield: Defending Against Backdoors on Text-to-Image Diffusion Models
by: Wang, Zhongqi, et al.
Published: (2024)
by: Wang, Zhongqi, et al.
Published: (2024)
Contrastive Learning of Person-independent Representations for Facial Action Unit Detection
by: Li, Yong, et al.
Published: (2024)
by: Li, Yong, et al.
Published: (2024)
Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness
by: Wang, Sibo, et al.
Published: (2024)
by: Wang, Sibo, et al.
Published: (2024)
Anonymization Prompt Learning for Facial Privacy-Preserving Text-to-Image Generation
by: Shi, Liang, et al.
Published: (2024)
by: Shi, Liang, et al.
Published: (2024)
Confidence Aware Learning for Reliable Face Anti-spoofing
by: Long, Xingming, et al.
Published: (2024)
by: Long, Xingming, et al.
Published: (2024)
Revisiting Face Forgery Detection: From Facial Representation to Forgery Detection
by: Guo, Zonghui, et al.
Published: (2024)
by: Guo, Zonghui, et al.
Published: (2024)
V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs
by: Nie, Sen, et al.
Published: (2025)
by: Nie, Sen, et al.
Published: (2025)
Contrastive Spectral Rectification: Test-Time Defense towards Zero-shot Adversarial Robustness of CLIP
by: Nie, Sen, et al.
Published: (2026)
by: Nie, Sen, et al.
Published: (2026)
Towards Robust Semantic Segmentation against Patch-based Attack via Attention Refinement
by: Yuan, Zheng, et al.
Published: (2024)
by: Yuan, Zheng, et al.
Published: (2024)
Hallucination of Multimodal Large Language Models: A Survey
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
Trigger without Trace: Towards Stealthy Backdoor Attack on Text-to-Image Diffusion Models
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
FullLoRA: Efficiently Boosting the Robustness of Pretrained Vision Transformers
by: Yuan, Zheng, et al.
Published: (2024)
by: Yuan, Zheng, et al.
Published: (2024)
A Trustworthy Method for Multimodal Emotion Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
Steering Vision-Language Pre-trained Models for Incremental Face Presentation Attack Detection
by: Li, Haoze, et al.
Published: (2025)
by: Li, Haoze, et al.
Published: (2025)
Image to Pseudo-Episode: Boosting Few-Shot Segmentation by Unlabeled Data
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
GLip: A Global-Local Integrated Progressive Framework for Robust Visual Speech Recognition
by: Wang, Tianyue, et al.
Published: (2025)
by: Wang, Tianyue, et al.
Published: (2025)
BIMM: Brain Inspired Masked Modeling for Video Representation Learning
by: Wan, Zhifan, et al.
Published: (2024)
by: Wan, Zhifan, et al.
Published: (2024)
What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models
by: Nie, Sen, et al.
Published: (2026)
by: Nie, Sen, et al.
Published: (2026)
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
by: Wang, Sibo, et al.
Published: (2024)
by: Wang, Sibo, et al.
Published: (2024)
Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating
by: Cao, Xiangkui, et al.
Published: (2026)
by: Cao, Xiangkui, et al.
Published: (2026)
Component-Based Out-of-Distribution Detection
by: Liu, Wenrui, et al.
Published: (2026)
by: Liu, Wenrui, et al.
Published: (2026)
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
by: Hou, Ruibing, et al.
Published: (2025)
by: Hou, Ruibing, et al.
Published: (2025)
Segment Anything for Videos: A Systematic Survey
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
by: Zhao, Jiahe, et al.
Published: (2025)
by: Zhao, Jiahe, et al.
Published: (2025)
Video Unsupervised Domain Adaptation with Deep Learning: A Comprehensive Survey
by: Xu, Yuecong, et al.
Published: (2022)
by: Xu, Yuecong, et al.
Published: (2022)
Similar Items
-
SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
by: Yan, Bei, et al.
Published: (2025) -
ACT Now: Preempting LVLM Hallucinations via Adaptive Context Integration
by: Yan, Bei, et al.
Published: (2026) -
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
by: Yang, Junqi, et al.
Published: (2026) -
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
by: Yan, Bei, et al.
Published: (2024) -
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
by: Yan, Bei, et al.
Published: (2024)