AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gong, Kaixiong, Feng, Kaituo, Li, Bohao, Wang, Yibing, Cheng, Mofan, Yang, Shijia, Han, Jiaming, Wang, Benyou, Bai, Yutong, Yang, Zhuoran, Yue, Xiangyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910725806489600
author Gong, Kaixiong
Feng, Kaituo
Li, Bohao
Wang, Yibing
Cheng, Mofan
Yang, Shijia
Han, Jiaming
Wang, Benyou
Bai, Yutong
Yang, Zhuoran
Yue, Xiangyu
author_facet Gong, Kaixiong
Feng, Kaituo
Li, Bohao
Wang, Yibing
Cheng, Mofan
Yang, Shijia
Han, Jiaming
Wang, Benyou
Bai, Yutong
Yang, Zhuoran
Yue, Xiangyu
contents Recently, multimodal large language models (MLLMs), such as GPT-4o, Gemini 1.5 Pro, and Reka Core, have expanded their capabilities to include vision and audio modalities. While these models demonstrate impressive performance across a wide range of audio-visual applications, our proposed DeafTest reveals that MLLMs often struggle with simple tasks humans find trivial: 1) determining which of two sounds is louder, and 2) determining which of two sounds has a higher pitch. Motivated by these observations, we introduce AV-Odyssey Bench, a comprehensive audio-visual benchmark designed to assess whether those MLLMs can truly understand the audio-visual information. This benchmark encompasses 4,555 carefully crafted problems, each incorporating text, visual, and audio components. To successfully infer answers, models must effectively leverage clues from both visual and audio inputs. To ensure precise and objective evaluation of MLLM responses, we have structured the questions as multiple-choice, eliminating the need for human evaluation or LLM-assisted assessment. We benchmark a series of closed-source and open-source models and summarize the observations. By revealing the limitations of current models, we aim to provide useful insight for future dataset collection and model development.
format Preprint
id arxiv_https___arxiv_org_abs_2412_02611
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Gong, Kaixiong
Feng, Kaituo
Li, Bohao
Wang, Yibing
Cheng, Mofan
Yang, Shijia
Han, Jiaming
Wang, Benyou
Bai, Yutong
Yang, Zhuoran
Yue, Xiangyu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
Sound
Audio and Speech Processing
Recently, multimodal large language models (MLLMs), such as GPT-4o, Gemini 1.5 Pro, and Reka Core, have expanded their capabilities to include vision and audio modalities. While these models demonstrate impressive performance across a wide range of audio-visual applications, our proposed DeafTest reveals that MLLMs often struggle with simple tasks humans find trivial: 1) determining which of two sounds is louder, and 2) determining which of two sounds has a higher pitch. Motivated by these observations, we introduce AV-Odyssey Bench, a comprehensive audio-visual benchmark designed to assess whether those MLLMs can truly understand the audio-visual information. This benchmark encompasses 4,555 carefully crafted problems, each incorporating text, visual, and audio components. To successfully infer answers, models must effectively leverage clues from both visual and audio inputs. To ensure precise and objective evaluation of MLLM responses, we have structured the questions as multiple-choice, eliminating the need for human evaluation or LLM-assisted assessment. We benchmark a series of closed-source and open-source models and summarize the observations. By revealing the limitations of current models, we aim to provide useful insight for future dataset collection and model development.
title AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.02611