Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Hongcheng, Qu, Jiashu, Tang, Jingyi, Bi, Baolong, Liu, Yue, Chen, Hongyu, Liang, Li, Su, Li, Huang, Qingming |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation
di: Dang, Tiantian, et al.
Pubblicazione: (2026)
di: Dang, Tiantian, et al.
Pubblicazione: (2026)
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
di: Li, Jiale, et al.
Pubblicazione: (2025)
di: Li, Jiale, et al.
Pubblicazione: (2025)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
di: Li, Chaoyu, et al.
Pubblicazione: (2024)
di: Li, Chaoyu, et al.
Pubblicazione: (2024)
Treble Counterfactual VLMs: A Causal Approach to Hallucination
di: Li, Shawn, et al.
Pubblicazione: (2025)
di: Li, Shawn, et al.
Pubblicazione: (2025)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation
di: Pan, Jiadong, et al.
Pubblicazione: (2024)
di: Pan, Jiadong, et al.
Pubblicazione: (2024)
Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
di: Yue, Zihao, et al.
Pubblicazione: (2024)
di: Yue, Zihao, et al.
Pubblicazione: (2024)
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
di: Ma, Ji, et al.
Pubblicazione: (2026)
di: Ma, Ji, et al.
Pubblicazione: (2026)
Hallucination Mitigation Prompts Long-term Video Understanding
di: Sun, Yiwei, et al.
Pubblicazione: (2024)
di: Sun, Yiwei, et al.
Pubblicazione: (2024)
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
di: Shu, Yan, et al.
Pubblicazione: (2025)
di: Shu, Yan, et al.
Pubblicazione: (2025)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
di: Li, Zongxia, et al.
Pubblicazione: (2025)
di: Li, Zongxia, et al.
Pubblicazione: (2025)
Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models
di: Pan, Jiadong, et al.
Pubblicazione: (2026)
di: Pan, Jiadong, et al.
Pubblicazione: (2026)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
di: Zheng, Kening, et al.
Pubblicazione: (2024)
di: Zheng, Kening, et al.
Pubblicazione: (2024)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
di: Seth, Ashish, et al.
Pubblicazione: (2025)
di: Seth, Ashish, et al.
Pubblicazione: (2025)
Exploring Audio Hallucination in Egocentric Video Understanding
di: Seth, Ashish, et al.
Pubblicazione: (2026)
di: Seth, Ashish, et al.
Pubblicazione: (2026)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
di: He, Zhentao, et al.
Pubblicazione: (2025)
di: He, Zhentao, et al.
Pubblicazione: (2025)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models
di: Rawte, Vipula, et al.
Pubblicazione: (2024)
di: Rawte, Vipula, et al.
Pubblicazione: (2024)
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
di: Lu, Hao, et al.
Pubblicazione: (2025)
di: Lu, Hao, et al.
Pubblicazione: (2025)
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
di: Sun, Yiming, et al.
Pubblicazione: (2025)
di: Sun, Yiming, et al.
Pubblicazione: (2025)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
di: Wang, Youze, et al.
Pubblicazione: (2025)
di: Wang, Youze, et al.
Pubblicazione: (2025)
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
di: Lu, Yifan, et al.
Pubblicazione: (2025)
di: Lu, Yifan, et al.
Pubblicazione: (2025)
Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models
di: Sun, Li, et al.
Pubblicazione: (2024)
di: Sun, Li, et al.
Pubblicazione: (2024)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
Exploring Causes and Mitigation of Hallucinations in Large Vision Language Models
di: Sun, Yaqi, et al.
Pubblicazione: (2025)
di: Sun, Yaqi, et al.
Pubblicazione: (2025)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
di: Zhu, Xingyu, et al.
Pubblicazione: (2026)
di: Zhu, Xingyu, et al.
Pubblicazione: (2026)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
di: Li, Yun, et al.
Pubblicazione: (2025)
di: Li, Yun, et al.
Pubblicazione: (2025)
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
di: Yang, Garry, et al.
Pubblicazione: (2025)
di: Yang, Garry, et al.
Pubblicazione: (2025)
Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models
di: Xing, Wenbin, et al.
Pubblicazione: (2026)
di: Xing, Wenbin, et al.
Pubblicazione: (2026)
Mitigating Multilingual Hallucination in Large Vision-Language Models
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
di: Fan, Linfeng, et al.
Pubblicazione: (2026)
di: Fan, Linfeng, et al.
Pubblicazione: (2026)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
di: Chang, Yue, et al.
Pubblicazione: (2024)
di: Chang, Yue, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Video Large Language Models via Spatiotemporal-Semantic Contrastive Decoding
di: Gao, Yuansheng, et al.
Pubblicazione: (2026)
di: Gao, Yuansheng, et al.
Pubblicazione: (2026)
Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
di: Tang, Kai, et al.
Pubblicazione: (2025)
di: Tang, Kai, et al.
Pubblicazione: (2025)
Identify, Isolate, and Purge: Mitigating Hallucinations in LVLMs via Self-Evolving Distillation
di: Li, Wenhao, et al.
Pubblicazione: (2025)
di: Li, Wenhao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation
di: Dang, Tiantian, et al.
Pubblicazione: (2026) -
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
di: Li, Jiale, et al.
Pubblicazione: (2025) -
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
di: Li, Chaoyu, et al.
Pubblicazione: (2024) -
Treble Counterfactual VLMs: A Causal Approach to Hallucination
di: Li, Shawn, et al.
Pubblicazione: (2025) -
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
di: Zhong, Weihong, et al.
Pubblicazione: (2024)