MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sakshi, S, Tyagi, Utkarsh, Kumar, Sonal, Seth, Ashish, Selvakumar, Ramaneswaran, Nieto, Oriol, Duraiswami, Ramani, Ghosh, Sreyan, Manocha, Dinesh
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917815841193984
author Sakshi, S
Tyagi, Utkarsh
Kumar, Sonal
Seth, Ashish
Selvakumar, Ramaneswaran
Nieto, Oriol
Duraiswami, Ramani
Ghosh, Sreyan
Manocha, Dinesh
author_facet Sakshi, S
Tyagi, Utkarsh
Kumar, Sonal
Seth, Ashish
Selvakumar, Ramaneswaran
Nieto, Oriol
Duraiswami, Ramani
Ghosh, Sreyan
Manocha, Dinesh
contents The ability to comprehend audio--which includes speech, non-speech sounds, and music--is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex reasoning. MMAU comprises 10k carefully curated audio clips paired with human-annotated natural language questions and answers spanning speech, environmental sounds, and music. It includes information extraction and reasoning questions, requiring models to demonstrate 27 distinct skills across unique and challenging tasks. Unlike existing benchmarks, MMAU emphasizes advanced perception and reasoning with domain-specific knowledge, challenging models to tackle tasks akin to those faced by experts. We assess 18 open-source and proprietary (Large) Audio-Language Models, demonstrating the significant challenges posed by MMAU. Notably, even the most advanced Gemini Pro v1.5 achieves only 52.97% accuracy, and the state-of-the-art open-source Qwen2-Audio achieves only 52.50%, highlighting considerable room for improvement. We believe MMAU will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19168
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
Sakshi, S
Tyagi, Utkarsh
Kumar, Sonal
Seth, Ashish
Selvakumar, Ramaneswaran
Nieto, Oriol
Duraiswami, Ramani
Ghosh, Sreyan
Manocha, Dinesh
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
The ability to comprehend audio--which includes speech, non-speech sounds, and music--is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex reasoning. MMAU comprises 10k carefully curated audio clips paired with human-annotated natural language questions and answers spanning speech, environmental sounds, and music. It includes information extraction and reasoning questions, requiring models to demonstrate 27 distinct skills across unique and challenging tasks. Unlike existing benchmarks, MMAU emphasizes advanced perception and reasoning with domain-specific knowledge, challenging models to tackle tasks akin to those faced by experts. We assess 18 open-source and proprietary (Large) Audio-Language Models, demonstrating the significant challenges posed by MMAU. Notably, even the most advanced Gemini Pro v1.5 achieves only 52.97% accuracy, and the state-of-the-art open-source Qwen2-Audio achieves only 52.50%, highlighting considerable room for improvement. We believe MMAU will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.
title MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
url https://arxiv.org/abs/2410.19168