AudioBench: A Universal Benchmark for Audio Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913821595009024 |
|---|---|
| author | Wang, Bin Zou, Xunlong Lin, Geyu Sun, Shuo Liu, Zhuohan Zhang, Wenyu Liu, Zhengyuan Aw, AiTi Chen, Nancy F. |
| author_facet | Wang, Bin Zou, Xunlong Lin, Geyu Sun, Shuo Liu, Zhuohan Zhang, Wenyu Liu, Zhengyuan Aw, AiTi Chen, Nancy F. |
| contents | We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main aspects: speech understanding, audio scene understanding, and voice understanding (paralinguistic). Despite recent advancements, there lacks a comprehensive benchmark for AudioLLMs on instruction following capabilities conditioned on audio signals. AudioBench addresses this gap by setting up datasets as well as desired evaluation metrics. Besides, we also evaluated the capabilities of five popular models and found that no single model excels consistently across all tasks. We outline the research outlook for AudioLLMs and anticipate that our open-sourced evaluation toolkit, data, and leaderboard will offer a robust testbed for future model developments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_16020 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | AudioBench: A Universal Benchmark for Audio Large Language Models Wang, Bin Zou, Xunlong Lin, Geyu Sun, Shuo Liu, Zhuohan Zhang, Wenyu Liu, Zhengyuan Aw, AiTi Chen, Nancy F. Sound Computation and Language Audio and Speech Processing We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main aspects: speech understanding, audio scene understanding, and voice understanding (paralinguistic). Despite recent advancements, there lacks a comprehensive benchmark for AudioLLMs on instruction following capabilities conditioned on audio signals. AudioBench addresses this gap by setting up datasets as well as desired evaluation metrics. Besides, we also evaluated the capabilities of five popular models and found that no single model excels consistently across all tasks. We outline the research outlook for AudioLLMs and anticipate that our open-sourced evaluation toolkit, data, and leaderboard will offer a robust testbed for future model developments. |
| title | AudioBench: A Universal Benchmark for Audio Large Language Models |
| topic | Sound Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2406.16020 |