The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
Fuente:
arXiv
Saved in:
| Main Authors: | Leng, Sicong, Xing, Yun, Cheng, Zesen, Zhou, Yang, Zhang, Hang, Li, Xin, Zhao, Deli, Lu, Shijian, Miao, Chunyan, Bing, Lidong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
Breaking the Memory Barrier: Near Infinite Batch Size Scaling for Contrastive Loss
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
by: Zhang, Boqiang, et al.
Published: (2025)
by: Zhang, Boqiang, et al.
Published: (2025)
MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
by: Leng, Sicong, et al.
Published: (2025)
by: Leng, Sicong, et al.
Published: (2025)
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
by: Qiu, Han, et al.
Published: (2024)
by: Qiu, Han, et al.
Published: (2024)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
Mitigating Object Hallucination via Concentric Causal Attention
by: Xing, Yun, et al.
Published: (2024)
by: Xing, Yun, et al.
Published: (2024)
Referring Multiple Regions with Large Multimodal Models via Contextual Latent Steering
by: Xing, Yun, et al.
Published: (2026)
by: Xing, Yun, et al.
Published: (2026)
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
by: Zhang, Wenqi, et al.
Published: (2025)
by: Zhang, Wenqi, et al.
Published: (2025)
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
by: Nie, Jiahao, et al.
Published: (2024)
by: Nie, Jiahao, et al.
Published: (2024)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024)
by: Yuan, Yuqian, et al.
Published: (2024)
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
Disentangled Graph Variational Auto-Encoder for Multimodal Recommendation with Interpretability
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
by: Sung-Bin, Kim, et al.
Published: (2024)
by: Sung-Bin, Kim, et al.
Published: (2024)
Learning Visual Affordance from Audio
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
by: Xing, Bohao, et al.
Published: (2025)
by: Xing, Bohao, et al.
Published: (2025)
Segment Anything with Multiple Modalities
by: Xiao, Aoran, et al.
Published: (2024)
by: Xiao, Aoran, et al.
Published: (2024)
Research on the Driving Mechanism of Network Structure of Patent‐Intensive Industrial Clusters: Taking Biopharmaceutical Industry as an Example
by: Lanqing Ge, et al.
Published: (2025)
by: Lanqing Ge, et al.
Published: (2025)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
by: Lin, Meng-Ping, et al.
Published: (2025)
by: Lin, Meng-Ping, et al.
Published: (2025)
Benford's Curse: Tracing Digit Bias to Numerical Hallucination in LLMs
by: Shao, Jiandong, et al.
Published: (2025)
by: Shao, Jiandong, et al.
Published: (2025)
Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
by: Park, Hansol, et al.
Published: (2025)
by: Park, Hansol, et al.
Published: (2025)
Boosting Reasoning in Large Multimodal Models via Activation Replay
by: Xing, Yun, et al.
Published: (2025)
by: Xing, Yun, et al.
Published: (2025)
Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
by: Zhao, Ruochen, et al.
Published: (2024)
by: Zhao, Ruochen, et al.
Published: (2024)
Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
by: Hsu, Tzu-wen, et al.
Published: (2025)
by: Hsu, Tzu-wen, et al.
Published: (2025)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
by: Xie, Zhifei, et al.
Published: (2025)
by: Xie, Zhifei, et al.
Published: (2025)
Multi-Modal Hallucination Control by Visual Information Grounding
by: Favero, Alessandro, et al.
Published: (2024)
by: Favero, Alessandro, et al.
Published: (2024)
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
by: Pian, Weiguo, et al.
Published: (2024)
by: Pian, Weiguo, et al.
Published: (2024)
Audio-Visual Intelligence in Large Foundation Models
by: Qin, You, et al.
Published: (2026)
by: Qin, You, et al.
Published: (2026)
Reducing Hallucinations of Medical Multimodal Large Language Models with Visual Retrieval-Augmented Generation
by: Chu, Yun-Wei, et al.
Published: (2025)
by: Chu, Yun-Wei, et al.
Published: (2025)
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
by: Li, Duo, et al.
Published: (2025)
by: Li, Duo, et al.
Published: (2025)
Self-Adapting Large Visual-Language Models to Edge Devices across Visual Modalities
by: Cai, Kaiwen, et al.
Published: (2024)
by: Cai, Kaiwen, et al.
Published: (2024)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
by: Zhou, Ziwei, et al.
Published: (2025)
by: Zhou, Ziwei, et al.
Published: (2025)
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
by: Xu, Ke, et al.
Published: (2026)
by: Xu, Ke, et al.
Published: (2026)
MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models
by: Chung, Sangyun, et al.
Published: (2026)
by: Chung, Sangyun, et al.
Published: (2026)
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
by: Zhou, Guanyu, et al.
Published: (2024)
by: Zhou, Guanyu, et al.
Published: (2024)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
On the Audio Hallucinations in Large Audio-Video Language Models
by: Nishimura, Taichi, et al.
Published: (2024)
by: Nishimura, Taichi, et al.
Published: (2024)
Visual Object Tracking across Diverse Data Modalities: A Review
by: Wang, Mengmeng, et al.
Published: (2024)
by: Wang, Mengmeng, et al.
Published: (2024)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
by: Chia, Yew Ken, et al.
Published: (2024)
by: Chia, Yew Ken, et al.
Published: (2024)
Similar Items
-
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
by: Cheng, Zesen, et al.
Published: (2024) -
Breaking the Memory Barrier: Near Infinite Batch Size Scaling for Contrastive Loss
by: Cheng, Zesen, et al.
Published: (2024) -
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
by: Zhang, Boqiang, et al.
Published: (2025) -
MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
by: Leng, Sicong, et al.
Published: (2025) -
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
by: Qiu, Han, et al.
Published: (2024)