Distorted or Fabricated? A Survey on Hallucination in Video LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Yiyang, Zhang, Yitian, Wang, Yizhou, Zhang, Mingyuan, Shi, Liang, Zeng, Huimin, Fu, Yun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
von: Huang, Yiyang, et al.
Veröffentlicht: (2025)
von: Huang, Yiyang, et al.
Veröffentlicht: (2025)
D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
von: Huang, Yiyang, et al.
Veröffentlicht: (2025)
von: Huang, Yiyang, et al.
Veröffentlicht: (2025)
Don't Judge by the Look: Towards Motion Coherent Video Representation
von: Zhang, Yitian, et al.
Veröffentlicht: (2024)
von: Zhang, Yitian, et al.
Veröffentlicht: (2024)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models
von: Chen, Zhawnen, et al.
Veröffentlicht: (2024)
von: Chen, Zhawnen, et al.
Veröffentlicht: (2024)
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
von: Yang, Junqi, et al.
Veröffentlicht: (2026)
von: Yang, Junqi, et al.
Veröffentlicht: (2026)
Trajectory Prediction Meets Large Language Models: A Survey
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
Accessing Vision Foundation Models via ImageNet-1K
von: Zhang, Yitian, et al.
Veröffentlicht: (2024)
von: Zhang, Yitian, et al.
Veröffentlicht: (2024)
Intuitive Axial Augmentation Using Polar-Sine-Based Piecewise Distortion for Medical Slice-Wise Segmentation
von: Zhang, Yiqin, et al.
Veröffentlicht: (2024)
von: Zhang, Yiqin, et al.
Veröffentlicht: (2024)
REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
von: Zhang, Yitian, et al.
Veröffentlicht: (2025)
von: Zhang, Yitian, et al.
Veröffentlicht: (2025)
SCBench: A Sports Commentary Benchmark for Video LLMs
von: Ge, Kuangzhi, et al.
Veröffentlicht: (2024)
von: Ge, Kuangzhi, et al.
Veröffentlicht: (2024)
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
von: Lu, Hao, et al.
Veröffentlicht: (2025)
von: Lu, Hao, et al.
Veröffentlicht: (2025)
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
von: Chen, Cong, et al.
Veröffentlicht: (2025)
von: Chen, Cong, et al.
Veröffentlicht: (2025)
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
von: Chen, Yuqing, et al.
Veröffentlicht: (2025)
von: Chen, Yuqing, et al.
Veröffentlicht: (2025)
Exploring Audio Hallucination in Egocentric Video Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
DAOVI: Distortion-Aware Omnidirectional Video Inpainting
von: Seshimo, Ryosuke, et al.
Veröffentlicht: (2025)
von: Seshimo, Ryosuke, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Video Large Language Models via Spatiotemporal-Semantic Contrastive Decoding
von: Gao, Yuansheng, et al.
Veröffentlicht: (2026)
von: Gao, Yuansheng, et al.
Veröffentlicht: (2026)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
Segment Anything for Videos: A Systematic Survey
von: Zhang, Chunhui, et al.
Veröffentlicht: (2024)
von: Zhang, Chunhui, et al.
Veröffentlicht: (2024)
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs
von: Lin, Chenchen, et al.
Veröffentlicht: (2026)
von: Lin, Chenchen, et al.
Veröffentlicht: (2026)
VERHallu: Evaluating and Mitigating Event Relation Hallucination in Video Large Language Models
von: Zhang, Zefan, et al.
Veröffentlicht: (2026)
von: Zhang, Zefan, et al.
Veröffentlicht: (2026)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
von: Peng, Tianhao, et al.
Veröffentlicht: (2025)
von: Peng, Tianhao, et al.
Veröffentlicht: (2025)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2025)
von: Seth, Ashish, et al.
Veröffentlicht: (2025)
CHASD: Language Increment-Calibrated Contrastive Decoding against Hallucination in LVLMs
von: Huang, Xiaoyi, et al.
Veröffentlicht: (2026)
von: Huang, Xiaoyi, et al.
Veröffentlicht: (2026)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
VideoQA in the Era of LLMs: An Empirical Study
von: Xiao, Junbin, et al.
Veröffentlicht: (2024)
von: Xiao, Junbin, et al.
Veröffentlicht: (2024)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
von: Pan, Yaning, et al.
Veröffentlicht: (2025)
von: Pan, Yaning, et al.
Veröffentlicht: (2025)
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
von: Yang, Garry, et al.
Veröffentlicht: (2025)
von: Yang, Garry, et al.
Veröffentlicht: (2025)
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
von: Zhang, Yuanhong, et al.
Veröffentlicht: (2026)
von: Zhang, Yuanhong, et al.
Veröffentlicht: (2026)
SAM2 for Image and Video Segmentation: A Comprehensive Survey
von: Jiaxing, Zhang, et al.
Veröffentlicht: (2025)
von: Jiaxing, Zhang, et al.
Veröffentlicht: (2025)
TPC: Cross-Temporal Prediction Connection for Vision-Language Model Hallucination Reduction
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
Synergistic Global-space Camera and Human Reconstruction from Videos
von: Zhao, Yizhou, et al.
Veröffentlicht: (2024)
von: Zhao, Yizhou, et al.
Veröffentlicht: (2024)
Hallucination-aware intermediate representation edit in large vision-language models
von: Suo, Wei, et al.
Veröffentlicht: (2026)
von: Suo, Wei, et al.
Veröffentlicht: (2026)
Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding
von: Suo, Wei, et al.
Veröffentlicht: (2025)
von: Suo, Wei, et al.
Veröffentlicht: (2025)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
FlowAct-R1: Towards Interactive Humanoid Video Generation
von: Wang, Lizhen, et al.
Veröffentlicht: (2026)
von: Wang, Lizhen, et al.
Veröffentlicht: (2026)
Memory-based Cross-modal Semantic Alignment Network for Radiology Report Generation
von: Tao, Yitian, et al.
Veröffentlicht: (2024)
von: Tao, Yitian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
von: Huang, Yiyang, et al.
Veröffentlicht: (2025) -
D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
von: Huang, Yiyang, et al.
Veröffentlicht: (2025) -
Don't Judge by the Look: Towards Motion Coherent Video Representation
von: Zhang, Yitian, et al.
Veröffentlicht: (2024) -
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
von: Dong, Qihua, et al.
Veröffentlicht: (2026) -
Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models
von: Chen, Zhawnen, et al.
Veröffentlicht: (2024)