MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Garry, Chen, Zizhe, Wong, Man Hon, Lei, Haoyu, Chen, Yongqiang, Li, Zhenguo, Zhou, Kaiwen, Cheng, James |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
par: Gao, Hongcheng, et autres
Publié: (2025)
par: Gao, Hongcheng, et autres
Publié: (2025)
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
par: Zhu, Lei, et autres
Publié: (2026)
par: Zhu, Lei, et autres
Publié: (2026)
MotionLLM: Understanding Human Behaviors from Human Motions and Videos
par: Chen, Ling-Hao, et autres
Publié: (2024)
par: Chen, Ling-Hao, et autres
Publié: (2024)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
par: Li, Zongxia, et autres
Publié: (2025)
par: Li, Zongxia, et autres
Publié: (2025)
Hallucination Mitigation Prompts Long-term Video Understanding
par: Sun, Yiwei, et autres
Publié: (2024)
par: Sun, Yiwei, et autres
Publié: (2024)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
par: Zhao, Jiaxing, et autres
Publié: (2025)
par: Zhao, Jiaxing, et autres
Publié: (2025)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
par: Li, Chaoyu, et autres
Publié: (2024)
par: Li, Chaoyu, et autres
Publié: (2024)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
par: Seth, Ashish, et autres
Publié: (2025)
par: Seth, Ashish, et autres
Publié: (2025)
Exploring Audio Hallucination in Egocentric Video Understanding
par: Seth, Ashish, et autres
Publié: (2026)
par: Seth, Ashish, et autres
Publié: (2026)
Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
par: Xu, Lu, et autres
Publié: (2024)
par: Xu, Lu, et autres
Publié: (2024)
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
par: Zhang, Jiacheng, et autres
Publié: (2024)
par: Zhang, Jiacheng, et autres
Publié: (2024)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
par: Yin, Yufei, et autres
Publié: (2026)
par: Yin, Yufei, et autres
Publié: (2026)
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
par: Sun, Yiming, et autres
Publié: (2025)
par: Sun, Yiming, et autres
Publié: (2025)
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
par: Pu, Bowei, et autres
Publié: (2025)
par: Pu, Bowei, et autres
Publié: (2025)
A Unified Perspective on Adversarial Membership Manipulation in Vision Models
par: Gao, Ruize, et autres
Publié: (2026)
par: Gao, Ruize, et autres
Publié: (2026)
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
par: Wang, Yuxuan, et autres
Publié: (2024)
par: Wang, Yuxuan, et autres
Publié: (2024)
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
par: Huang, Zhe, et autres
Publié: (2025)
par: Huang, Zhe, et autres
Publié: (2025)
LongVLM: Efficient Long Video Understanding via Large Language Models
par: Weng, Yuetian, et autres
Publié: (2024)
par: Weng, Yuetian, et autres
Publié: (2024)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
par: Zhi, Zhuo, et autres
Publié: (2025)
par: Zhi, Zhuo, et autres
Publié: (2025)
VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
par: Cheng, Ying, et autres
Publié: (2025)
par: Cheng, Ying, et autres
Publié: (2025)
EEA: Exploration-Exploitation Agent for Long Video Understanding
par: Yang, Te, et autres
Publié: (2025)
par: Yang, Te, et autres
Publié: (2025)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
par: Chen, Yuxiao, et autres
Publié: (2026)
par: Chen, Yuxiao, et autres
Publié: (2026)
TempCompass: Do Video LLMs Really Understand Videos?
par: Liu, Yuanxin, et autres
Publié: (2024)
par: Liu, Yuanxin, et autres
Publié: (2024)
Vidi: Large Multimodal Models for Video Understanding and Editing
par: Vidi Team, et autres
Publié: (2025)
par: Vidi Team, et autres
Publié: (2025)
Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning
par: Wang, Shida, et autres
Publié: (2026)
par: Wang, Shida, et autres
Publié: (2026)
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
par: Zhou, Ting, et autres
Publié: (2024)
par: Zhou, Ting, et autres
Publié: (2024)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
par: Chen, Shimin, et autres
Publié: (2024)
par: Chen, Shimin, et autres
Publié: (2024)
VCA: Video Curious Agent for Long Video Understanding
par: Yang, Zeyuan, et autres
Publié: (2024)
par: Yang, Zeyuan, et autres
Publié: (2024)
EventRR: Event Referential Reasoning for Referring Video Object Segmentation
par: Xu, Huihui, et autres
Publié: (2025)
par: Xu, Huihui, et autres
Publié: (2025)
Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
par: Jiang, Xixi, et autres
Publié: (2025)
par: Jiang, Xixi, et autres
Publié: (2025)
Large Motion Video Autoencoding with Cross-modal Video VAE
par: Xing, Yazhou, et autres
Publié: (2024)
par: Xing, Yazhou, et autres
Publié: (2024)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
par: Cai, Yuxuan, et autres
Publié: (2025)
par: Cai, Yuxuan, et autres
Publié: (2025)
Personalized Video Summarization by Multimodal Video Understanding
par: Chen, Brian, et autres
Publié: (2024)
par: Chen, Brian, et autres
Publié: (2024)
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
par: Chu, Zhixuan, et autres
Publié: (2024)
par: Chu, Zhixuan, et autres
Publié: (2024)
Towards Event-oriented Long Video Understanding
par: Du, Yifan, et autres
Publié: (2024)
par: Du, Yifan, et autres
Publié: (2024)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
par: Zhang, Zicheng, et autres
Publié: (2024)
par: Zhang, Zicheng, et autres
Publié: (2024)
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
par: Ren, Weiming, et autres
Publié: (2024)
par: Ren, Weiming, et autres
Publié: (2024)
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
par: Cheng, Junhao, et autres
Publié: (2025)
par: Cheng, Junhao, et autres
Publié: (2025)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
par: Wang, Zhao, et autres
Publié: (2024)
par: Wang, Zhao, et autres
Publié: (2024)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
par: Chen, Houlun, et autres
Publié: (2026)
par: Chen, Houlun, et autres
Publié: (2026)
Documents similaires
-
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
par: Gao, Hongcheng, et autres
Publié: (2025) -
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
par: Zhu, Lei, et autres
Publié: (2026) -
MotionLLM: Understanding Human Behaviors from Human Motions and Videos
par: Chen, Ling-Hao, et autres
Publié: (2024) -
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
par: Li, Zongxia, et autres
Publié: (2025) -
Hallucination Mitigation Prompts Long-term Video Understanding
par: Sun, Yiwei, et autres
Publié: (2024)