Saved in:
Bibliographic Details
Main Authors: Chao, Jianghan, Gao, Jianzhang, Tan, Wenhui, Sun, Yuchong, Song, Ruihua, Ru, Liyun
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.12772
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910220205162496
author Chao, Jianghan
Gao, Jianzhang
Tan, Wenhui
Sun, Yuchong
Song, Ruihua
Ru, Liyun
author_facet Chao, Jianghan
Gao, Jianzhang
Tan, Wenhui
Sun, Yuchong
Song, Ruihua
Ru, Liyun
contents Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including vision and audio, an effective benchmark must comprehensively cover three key aspects: (1) multi-modal dependency (i.e., questions that cannot be answered using vision or audio alone), (2) diverse audio information types (e.g., speech, sound events), and (3) varying scene spans. However, existing datasets fall short in one or more of these dimensions, limiting strict and comprehensive evaluation. To address this gap, we introduce JointAVBench, a novel benchmark with strict audio-video correlation, spanning five cognitive dimensions, four audio information types (speech, sound events, music, vocal traits), and three scene spans (single-, cross-, and full-scene). Given the high cost of manual annotation, we propose an automated pipeline that leverages state-of-the-art vision-LLMs, audio-LLMs, and general-purpose LLMs to synthesize questions and answers that strictly require joint audio-visual understanding. We evaluate leading vision-only, audio-only, and Omni-LLMs on our dataset. Results show that even the best-performing Omni-LLM achieves an average accuracy of only 65.3\%, outperforming uni-modal baselines but revealing substantial room for improvement, especially in cross-scene reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12772
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
Chao, Jianghan
Gao, Jianzhang
Tan, Wenhui
Sun, Yuchong
Song, Ruihua
Ru, Liyun
Multimedia
Computer Vision and Pattern Recognition
Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including vision and audio, an effective benchmark must comprehensively cover three key aspects: (1) multi-modal dependency (i.e., questions that cannot be answered using vision or audio alone), (2) diverse audio information types (e.g., speech, sound events), and (3) varying scene spans. However, existing datasets fall short in one or more of these dimensions, limiting strict and comprehensive evaluation. To address this gap, we introduce JointAVBench, a novel benchmark with strict audio-video correlation, spanning five cognitive dimensions, four audio information types (speech, sound events, music, vocal traits), and three scene spans (single-, cross-, and full-scene). Given the high cost of manual annotation, we propose an automated pipeline that leverages state-of-the-art vision-LLMs, audio-LLMs, and general-purpose LLMs to synthesize questions and answers that strictly require joint audio-visual understanding. We evaluate leading vision-only, audio-only, and Omni-LLMs on our dataset. Results show that even the best-performing Omni-LLM achieves an average accuracy of only 65.3\%, outperforming uni-modal baselines but revealing substantial room for improvement, especially in cross-scene reasoning.
title JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
topic Multimedia
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.12772