MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Garry, Chen, Zizhe, Wong, Man Hon, Lei, Haoyu, Chen, Yongqiang, Li, Zhenguo, Zhou, Kaiwen, Cheng, James
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915489733672960
author Yang, Garry
Chen, Zizhe
Wong, Man Hon
Lei, Haoyu
Chen, Yongqiang
Li, Zhenguo
Zhou, Kaiwen
Cheng, James
author_facet Yang, Garry
Chen, Zizhe
Wong, Man Hon
Lei, Haoyu
Chen, Yongqiang
Li, Zhenguo
Zhou, Kaiwen
Cheng, James
contents Large Video Models (LVMs) build on the semantic capabilities of Large Language Models (LLMs) and vision modules by integrating temporal information to better understand dynamic video content. Despite their progress, LVMs are prone to hallucinations-producing inaccurate or irrelevant descriptions. Current benchmarks for video hallucination depend heavily on manual categorization of video content, neglecting the perception-based processes through which humans naturally interpret videos. We introduce MESH, a benchmark designed to evaluate hallucinations in LVMs systematically. MESH uses a Question-Answering framework with binary and multi-choice formats incorporating target and trap instances. It follows a bottom-up approach, evaluating basic objects, coarse-to-fine subject features, and subject-action pairs, aligning with human video understanding. We demonstrate that MESH offers an effective and comprehensive approach for identifying hallucinations in videos. Our evaluations show that while LVMs excel at recognizing basic objects and features, their susceptibility to hallucinations increases markedly when handling fine details or aligning multiple actions involving various subjects in longer videos.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08538
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
Yang, Garry
Chen, Zizhe
Wong, Man Hon
Lei, Haoyu
Chen, Yongqiang
Li, Zhenguo
Zhou, Kaiwen
Cheng, James
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Video Models (LVMs) build on the semantic capabilities of Large Language Models (LLMs) and vision modules by integrating temporal information to better understand dynamic video content. Despite their progress, LVMs are prone to hallucinations-producing inaccurate or irrelevant descriptions. Current benchmarks for video hallucination depend heavily on manual categorization of video content, neglecting the perception-based processes through which humans naturally interpret videos. We introduce MESH, a benchmark designed to evaluate hallucinations in LVMs systematically. MESH uses a Question-Answering framework with binary and multi-choice formats incorporating target and trap instances. It follows a bottom-up approach, evaluating basic objects, coarse-to-fine subject features, and subject-action pairs, aligning with human video understanding. We demonstrate that MESH offers an effective and comprehensive approach for identifying hallucinations in videos. Our evaluations show that while LVMs excel at recognizing basic objects and features, their susceptibility to hallucinations increases markedly when handling fine details or aligning multiple actions involving various subjects in longer videos.
title MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.08538