MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xie, Liuyue, Kuthiala, Avik, Wei, George Z., Zheng, Ce, Bal, Ananya, Dabhi, Mosam, Wen, Liting, Rustagi, Taru, Lai, Ethan, Khyalia, Sushil, Choudhury, Rohan, Ziyadi, Morteza, Zhang, Xu, Yang, Hao, Jeni, László A.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911305650143232
author Xie, Liuyue
Kuthiala, Avik
Wei, George Z.
Zheng, Ce
Bal, Ananya
Dabhi, Mosam
Wen, Liting
Rustagi, Taru
Lai, Ethan
Khyalia, Sushil
Choudhury, Rohan
Ziyadi, Morteza
Zhang, Xu
Yang, Hao
Jeni, László A.
author_facet Xie, Liuyue
Kuthiala, Avik
Wei, George Z.
Zheng, Ce
Bal, Ananya
Dabhi, Mosam
Wen, Liting
Rustagi, Taru
Lai, Ethan
Khyalia, Sushil
Choudhury, Rohan
Ziyadi, Morteza
Zhang, Xu
Yang, Hao
Jeni, László A.
contents We introduce MAVERIX (Multimodal audiovisual Evaluation and Recognition IndeX), a unified benchmark to probe the video understanding in multimodal LLMs, encompassing video, audio, text inputs with human performance baselines. Although recent advancements in models with vision and audio understanding capabilities have shown substantial progress, the field lacks a standardized evaluation framework to thoroughly assess their cross-modality comprehension performance. MAVERIX curates 2,556 questions from 700 videos, in the form of both multiple-choice and open-ended formats, explicitly designed to evaluate multimodal models through questions that necessitate tight integration of video and audio information, spanning a broad spectrum of agentic scenarios. MAVERIX uniquely provides models with audiovisual questions, closely mimicking the multimodal perceptual experiences available to humans during inference and decision-making processes. To our knowledge, MAVERIX is the first benchmark aimed explicitly at assessing comprehensive audiovisual integration in such granularity. Experiments with state-of-the-art models, including Qwen 2.5 Omni and Gemini 2.5 Flash-Lite, show performance around 64% accuracy, while human experts reach near-ceiling performance of 92.8%, exposing a substantial gap to human-level comprehension. With standardized evaluation protocols, a rigorously annotated pipeline, and a public toolkit, MAVERIX establishes a challenging testbed for advancing audiovisual multimodal intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21699
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX
Xie, Liuyue
Kuthiala, Avik
Wei, George Z.
Zheng, Ce
Bal, Ananya
Dabhi, Mosam
Wen, Liting
Rustagi, Taru
Lai, Ethan
Khyalia, Sushil
Choudhury, Rohan
Ziyadi, Morteza
Zhang, Xu
Yang, Hao
Jeni, László A.
Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
We introduce MAVERIX (Multimodal audiovisual Evaluation and Recognition IndeX), a unified benchmark to probe the video understanding in multimodal LLMs, encompassing video, audio, text inputs with human performance baselines. Although recent advancements in models with vision and audio understanding capabilities have shown substantial progress, the field lacks a standardized evaluation framework to thoroughly assess their cross-modality comprehension performance. MAVERIX curates 2,556 questions from 700 videos, in the form of both multiple-choice and open-ended formats, explicitly designed to evaluate multimodal models through questions that necessitate tight integration of video and audio information, spanning a broad spectrum of agentic scenarios. MAVERIX uniquely provides models with audiovisual questions, closely mimicking the multimodal perceptual experiences available to humans during inference and decision-making processes. To our knowledge, MAVERIX is the first benchmark aimed explicitly at assessing comprehensive audiovisual integration in such granularity. Experiments with state-of-the-art models, including Qwen 2.5 Omni and Gemini 2.5 Flash-Lite, show performance around 64% accuracy, while human experts reach near-ceiling performance of 92.8%, exposing a substantial gap to human-level comprehension. With standardized evaluation protocols, a rigorously annotated pipeline, and a public toolkit, MAVERIX establishes a challenging testbed for advancing audiovisual multimodal intelligence.
title MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX
topic Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2503.21699