OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908866365620224 |
|---|---|
| author | Li, Caorui Chen, Yu Ji, Yiyan Xu, Jin Cui, Zhenyu Li, Shihao Zhang, Yuanxing Wang, Wentao Song, Zhenghao Zhang, Dingling He, Ying Liu, Haoxiang Wang, Yuxuan Wang, Qiufeng Tang, Jiafu Wu, Zhenhe Luo, Jiehui Pan, Zhiyu Xie, Weihao Zhang, Chenchen Wang, Zhaohui Tian, Jiayi Wang, Yanghai Cao, Zhe Dai, Minxin Wang, Ke Wen, Runzhe Ma, Yinghao Pan, Yaning Chang, Sungkyun Taheri, Termeh Xia, Haiwen Plachouras, Christos Benetos, Emmanouil Li, Yizhi Zhang, Ge Yang, Jian Peng, Tianhao Wang, Zili Liu, Minghao Peng, Junran Zhang, Zhaoxiang Liu, Jiaheng |
| author_facet | Li, Caorui Chen, Yu Ji, Yiyan Xu, Jin Cui, Zhenyu Li, Shihao Zhang, Yuanxing Wang, Wentao Song, Zhenghao Zhang, Dingling He, Ying Liu, Haoxiang Wang, Yuxuan Wang, Qiufeng Tang, Jiafu Wu, Zhenhe Luo, Jiehui Pan, Zhiyu Xie, Weihao Zhang, Chenchen Wang, Zhaohui Tian, Jiayi Wang, Yanghai Cao, Zhe Dai, Minxin Wang, Ke Wen, Runzhe Ma, Yinghao Pan, Yaning Chang, Sungkyun Taheri, Termeh Xia, Haiwen Plachouras, Christos Benetos, Emmanouil Li, Yizhi Zhang, Ge Yang, Jian Peng, Tianhao Wang, Zili Liu, Minghao Peng, Junran Zhang, Zhaoxiang Liu, Jiaheng |
| contents | Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modalities or integrating them in a logically inconsistent manner. To bridge this gap, we introduce OmniVideoBench, a large-scale and rigorously designed benchmark dedicated to assessing synergistic audio-visual understanding, with a strong emphasis on modality complementarity and logical consistency. Specifically, OmniVideoBench comprises 1000 high-quality question-answer(QA) pairs, each annotated with step-by-step reasoning traces, derived from 628 diverse videos ranging from several seconds to 30 minutes, and manually verified to guarantee complete correctness and uniqueness. Moreover, OmniVideoBench encompasses 13 carefully designed question types, covering temporal reasoning, spatial localization, counting, causal inference, summarization, and beyond, thereby capturing the essential challenges of video understanding. Evaluation of multiple MLLMs on OmniVideoBench reveals a pronounced gap between model performance and human reasoning, with open-source models lagging significantly behind their closed-source counterparts, underscoring the inherent difficulty of genuine audio-visual reasoning. We will release OmniVideoBench to foster the development of MLLMs with stronger and more generalizable reasoning capabilities. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_10689 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs Li, Caorui Chen, Yu Ji, Yiyan Xu, Jin Cui, Zhenyu Li, Shihao Zhang, Yuanxing Wang, Wentao Song, Zhenghao Zhang, Dingling He, Ying Liu, Haoxiang Wang, Yuxuan Wang, Qiufeng Tang, Jiafu Wu, Zhenhe Luo, Jiehui Pan, Zhiyu Xie, Weihao Zhang, Chenchen Wang, Zhaohui Tian, Jiayi Wang, Yanghai Cao, Zhe Dai, Minxin Wang, Ke Wen, Runzhe Ma, Yinghao Pan, Yaning Chang, Sungkyun Taheri, Termeh Xia, Haiwen Plachouras, Christos Benetos, Emmanouil Li, Yizhi Zhang, Ge Yang, Jian Peng, Tianhao Wang, Zili Liu, Minghao Peng, Junran Zhang, Zhaoxiang Liu, Jiaheng Artificial Intelligence Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modalities or integrating them in a logically inconsistent manner. To bridge this gap, we introduce OmniVideoBench, a large-scale and rigorously designed benchmark dedicated to assessing synergistic audio-visual understanding, with a strong emphasis on modality complementarity and logical consistency. Specifically, OmniVideoBench comprises 1000 high-quality question-answer(QA) pairs, each annotated with step-by-step reasoning traces, derived from 628 diverse videos ranging from several seconds to 30 minutes, and manually verified to guarantee complete correctness and uniqueness. Moreover, OmniVideoBench encompasses 13 carefully designed question types, covering temporal reasoning, spatial localization, counting, causal inference, summarization, and beyond, thereby capturing the essential challenges of video understanding. Evaluation of multiple MLLMs on OmniVideoBench reveals a pronounced gap between model performance and human reasoning, with open-source models lagging significantly behind their closed-source counterparts, underscoring the inherent difficulty of genuine audio-visual reasoning. We will release OmniVideoBench to foster the development of MLLMs with stronger and more generalizable reasoning capabilities. |
| title | OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2510.10689 |