MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Tianhao, Wang, Haochen, Zhang, Yuanxing, Wang, Zekun, Wang, Zili, Chang, Gavin, Yang, Jian, Li, Shihao, Wang, Yanghai, Wang, Xintao, Li, Houyi, Ji, Wei, Wan, Pengfei, Huang, Steven, Zhang, Zhaoxiang, Liu, Jiaheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909900297207808
author Peng, Tianhao
Wang, Haochen
Zhang, Yuanxing
Wang, Zekun
Wang, Zili
Chang, Gavin
Yang, Jian
Li, Shihao
Wang, Yanghai
Wang, Xintao
Li, Houyi
Ji, Wei
Wan, Pengfei
Huang, Steven
Zhang, Zhaoxiang
Liu, Jiaheng
author_facet Peng, Tianhao
Wang, Haochen
Zhang, Yuanxing
Wang, Zekun
Wang, Zili
Chang, Gavin
Yang, Jian
Li, Shihao
Wang, Yanghai
Wang, Xintao
Li, Houyi
Ji, Wei
Wan, Pengfei
Huang, Steven
Zhang, Zhaoxiang
Liu, Jiaheng
contents The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark for evaluating Multi-Video Understanding for MLLMs. Specifically, our MVU-Eval mainly assesses eight core competencies through 1,824 meticulously curated question-answer pairs spanning 4,959 videos from diverse domains, addressing both fundamental perception tasks and high-order reasoning tasks. These capabilities are rigorously aligned with real-world applications such as multi-sensor synthesis in autonomous systems and cross-angle sports analytics. Through extensive evaluation of state-of-the-art open-source and closed-source models, we reveal significant performance discrepancies and limitations in current MLLMs' ability to perform understanding across multiple videos. The benchmark will be made publicly available to foster future research.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07250
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
Peng, Tianhao
Wang, Haochen
Zhang, Yuanxing
Wang, Zekun
Wang, Zili
Chang, Gavin
Yang, Jian
Li, Shihao
Wang, Yanghai
Wang, Xintao
Li, Houyi
Ji, Wei
Wan, Pengfei
Huang, Steven
Zhang, Zhaoxiang
Liu, Jiaheng
Computer Vision and Pattern Recognition
Artificial Intelligence
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark for evaluating Multi-Video Understanding for MLLMs. Specifically, our MVU-Eval mainly assesses eight core competencies through 1,824 meticulously curated question-answer pairs spanning 4,959 videos from diverse domains, addressing both fundamental perception tasks and high-order reasoning tasks. These capabilities are rigorously aligned with real-world applications such as multi-sensor synthesis in autonomous systems and cross-angle sports analytics. Through extensive evaluation of state-of-the-art open-source and closed-source models, we reveal significant performance discrepancies and limitations in current MLLMs' ability to perform understanding across multiple videos. The benchmark will be made publicly available to foster future research.
title MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.07250