OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Caorui, Chen, Yu, Ji, Yiyan, Xu, Jin, Cui, Zhenyu, Li, Shihao, Zhang, Yuanxing, Wang, Wentao, Song, Zhenghao, Zhang, Dingling, He, Ying, Liu, Haoxiang, Wang, Yuxuan, Wang, Qiufeng, Tang, Jiafu, Wu, Zhenhe, Luo, Jiehui, Pan, Zhiyu, Xie, Weihao, Zhang, Chenchen, Wang, Zhaohui, Tian, Jiayi, Wang, Yanghai, Cao, Zhe, Dai, Minxin, Wang, Ke, Wen, Runzhe, Ma, Yinghao, Pan, Yaning, Chang, Sungkyun, Taheri, Termeh, Xia, Haiwen, Plachouras, Christos, Benetos, Emmanouil, Li, Yizhi, Zhang, Ge, Yang, Jian, Peng, Tianhao, Wang, Zili, Liu, Minghao, Peng, Junran, Zhang, Zhaoxiang, Liu, Jiaheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908866365620224
author Li, Caorui
Chen, Yu
Ji, Yiyan
Xu, Jin
Cui, Zhenyu
Li, Shihao
Zhang, Yuanxing
Wang, Wentao
Song, Zhenghao
Zhang, Dingling
He, Ying
Liu, Haoxiang
Wang, Yuxuan
Wang, Qiufeng
Tang, Jiafu
Wu, Zhenhe
Luo, Jiehui
Pan, Zhiyu
Xie, Weihao
Zhang, Chenchen
Wang, Zhaohui
Tian, Jiayi
Wang, Yanghai
Cao, Zhe
Dai, Minxin
Wang, Ke
Wen, Runzhe
Ma, Yinghao
Pan, Yaning
Chang, Sungkyun
Taheri, Termeh
Xia, Haiwen
Plachouras, Christos
Benetos, Emmanouil
Li, Yizhi
Zhang, Ge
Yang, Jian
Peng, Tianhao
Wang, Zili
Liu, Minghao
Peng, Junran
Zhang, Zhaoxiang
Liu, Jiaheng
author_facet Li, Caorui
Chen, Yu
Ji, Yiyan
Xu, Jin
Cui, Zhenyu
Li, Shihao
Zhang, Yuanxing
Wang, Wentao
Song, Zhenghao
Zhang, Dingling
He, Ying
Liu, Haoxiang
Wang, Yuxuan
Wang, Qiufeng
Tang, Jiafu
Wu, Zhenhe
Luo, Jiehui
Pan, Zhiyu
Xie, Weihao
Zhang, Chenchen
Wang, Zhaohui
Tian, Jiayi
Wang, Yanghai
Cao, Zhe
Dai, Minxin
Wang, Ke
Wen, Runzhe
Ma, Yinghao
Pan, Yaning
Chang, Sungkyun
Taheri, Termeh
Xia, Haiwen
Plachouras, Christos
Benetos, Emmanouil
Li, Yizhi
Zhang, Ge
Yang, Jian
Peng, Tianhao
Wang, Zili
Liu, Minghao
Peng, Junran
Zhang, Zhaoxiang
Liu, Jiaheng
contents Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modalities or integrating them in a logically inconsistent manner. To bridge this gap, we introduce OmniVideoBench, a large-scale and rigorously designed benchmark dedicated to assessing synergistic audio-visual understanding, with a strong emphasis on modality complementarity and logical consistency. Specifically, OmniVideoBench comprises 1000 high-quality question-answer(QA) pairs, each annotated with step-by-step reasoning traces, derived from 628 diverse videos ranging from several seconds to 30 minutes, and manually verified to guarantee complete correctness and uniqueness. Moreover, OmniVideoBench encompasses 13 carefully designed question types, covering temporal reasoning, spatial localization, counting, causal inference, summarization, and beyond, thereby capturing the essential challenges of video understanding. Evaluation of multiple MLLMs on OmniVideoBench reveals a pronounced gap between model performance and human reasoning, with open-source models lagging significantly behind their closed-source counterparts, underscoring the inherent difficulty of genuine audio-visual reasoning. We will release OmniVideoBench to foster the development of MLLMs with stronger and more generalizable reasoning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10689
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
Li, Caorui
Chen, Yu
Ji, Yiyan
Xu, Jin
Cui, Zhenyu
Li, Shihao
Zhang, Yuanxing
Wang, Wentao
Song, Zhenghao
Zhang, Dingling
He, Ying
Liu, Haoxiang
Wang, Yuxuan
Wang, Qiufeng
Tang, Jiafu
Wu, Zhenhe
Luo, Jiehui
Pan, Zhiyu
Xie, Weihao
Zhang, Chenchen
Wang, Zhaohui
Tian, Jiayi
Wang, Yanghai
Cao, Zhe
Dai, Minxin
Wang, Ke
Wen, Runzhe
Ma, Yinghao
Pan, Yaning
Chang, Sungkyun
Taheri, Termeh
Xia, Haiwen
Plachouras, Christos
Benetos, Emmanouil
Li, Yizhi
Zhang, Ge
Yang, Jian
Peng, Tianhao
Wang, Zili
Liu, Minghao
Peng, Junran
Zhang, Zhaoxiang
Liu, Jiaheng
Artificial Intelligence
Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modalities or integrating them in a logically inconsistent manner. To bridge this gap, we introduce OmniVideoBench, a large-scale and rigorously designed benchmark dedicated to assessing synergistic audio-visual understanding, with a strong emphasis on modality complementarity and logical consistency. Specifically, OmniVideoBench comprises 1000 high-quality question-answer(QA) pairs, each annotated with step-by-step reasoning traces, derived from 628 diverse videos ranging from several seconds to 30 minutes, and manually verified to guarantee complete correctness and uniqueness. Moreover, OmniVideoBench encompasses 13 carefully designed question types, covering temporal reasoning, spatial localization, counting, causal inference, summarization, and beyond, thereby capturing the essential challenges of video understanding. Evaluation of multiple MLLMs on OmniVideoBench reveals a pronounced gap between model performance and human reasoning, with open-source models lagging significantly behind their closed-source counterparts, underscoring the inherent difficulty of genuine audio-visual reasoning. We will release OmniVideoBench to foster the development of MLLMs with stronger and more generalizable reasoning capabilities.
title OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
topic Artificial Intelligence
url https://arxiv.org/abs/2510.10689