AudioBench: A Universal Benchmark for Audio Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Bin, Zou, Xunlong, Lin, Geyu, Sun, Shuo, Liu, Zhuohan, Zhang, Wenyu, Liu, Zhengyuan, Aw, AiTi, Chen, Nancy F.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913821595009024
author Wang, Bin
Zou, Xunlong
Lin, Geyu
Sun, Shuo
Liu, Zhuohan
Zhang, Wenyu
Liu, Zhengyuan
Aw, AiTi
Chen, Nancy F.
author_facet Wang, Bin
Zou, Xunlong
Lin, Geyu
Sun, Shuo
Liu, Zhuohan
Zhang, Wenyu
Liu, Zhengyuan
Aw, AiTi
Chen, Nancy F.
contents We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main aspects: speech understanding, audio scene understanding, and voice understanding (paralinguistic). Despite recent advancements, there lacks a comprehensive benchmark for AudioLLMs on instruction following capabilities conditioned on audio signals. AudioBench addresses this gap by setting up datasets as well as desired evaluation metrics. Besides, we also evaluated the capabilities of five popular models and found that no single model excels consistently across all tasks. We outline the research outlook for AudioLLMs and anticipate that our open-sourced evaluation toolkit, data, and leaderboard will offer a robust testbed for future model developments.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16020
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AudioBench: A Universal Benchmark for Audio Large Language Models
Wang, Bin
Zou, Xunlong
Lin, Geyu
Sun, Shuo
Liu, Zhuohan
Zhang, Wenyu
Liu, Zhengyuan
Aw, AiTi
Chen, Nancy F.
Sound
Computation and Language
Audio and Speech Processing
We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main aspects: speech understanding, audio scene understanding, and voice understanding (paralinguistic). Despite recent advancements, there lacks a comprehensive benchmark for AudioLLMs on instruction following capabilities conditioned on audio signals. AudioBench addresses this gap by setting up datasets as well as desired evaluation metrics. Besides, we also evaluated the capabilities of five popular models and found that no single model excels consistently across all tasks. We outline the research outlook for AudioLLMs and anticipate that our open-sourced evaluation toolkit, data, and leaderboard will offer a robust testbed for future model developments.
title AudioBench: A Universal Benchmark for Audio Large Language Models
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2406.16020