The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiajia, Yang, Lu, Tang, Mingni, Chen, Cong, Li, Zuchao, Wang, Ping, Zhao, Hai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910499024666624
author Li, Jiajia
Yang, Lu
Tang, Mingni
Chen, Cong
Li, Zuchao
Wang, Ping
Zhao, Hai
author_facet Li, Jiajia
Yang, Lu
Tang, Mingni
Chen, Cong
Li, Zuchao
Wang, Ping
Zhao, Hai
contents Benchmark plays a pivotal role in assessing the advancements of large language models (LLMs). While numerous benchmarks have been proposed to evaluate LLMs' capabilities, there is a notable absence of a dedicated benchmark for assessing their musical abilities. To address this gap, we present ZIQI-Eval, a comprehensive and large-scale music benchmark specifically designed to evaluate the music-related capabilities of LLMs. ZIQI-Eval encompasses a wide range of questions, covering 10 major categories and 56 subcategories, resulting in over 14,000 meticulously curated data entries. By leveraging ZIQI-Eval, we conduct a comprehensive evaluation over 16 LLMs to evaluate and analyze LLMs' performance in the domain of music. Results indicate that all LLMs perform poorly on the ZIQI-Eval benchmark, suggesting significant room for improvement in their musical capabilities. With ZIQI-Eval, we aim to provide a standardized and robust evaluation framework that facilitates a comprehensive assessment of LLMs' music-related abilities. The dataset is available at GitHub\footnote{https://github.com/zcli-charlie/ZIQI-Eval} and HuggingFace\footnote{https://huggingface.co/datasets/MYTH-Lab/ZIQI-Eval}.
format Preprint
id arxiv_https___arxiv_org_abs_2406_15885
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models
Li, Jiajia
Yang, Lu
Tang, Mingni
Chen, Cong
Li, Zuchao
Wang, Ping
Zhao, Hai
Sound
Artificial Intelligence
Audio and Speech Processing
Benchmark plays a pivotal role in assessing the advancements of large language models (LLMs). While numerous benchmarks have been proposed to evaluate LLMs' capabilities, there is a notable absence of a dedicated benchmark for assessing their musical abilities. To address this gap, we present ZIQI-Eval, a comprehensive and large-scale music benchmark specifically designed to evaluate the music-related capabilities of LLMs. ZIQI-Eval encompasses a wide range of questions, covering 10 major categories and 56 subcategories, resulting in over 14,000 meticulously curated data entries. By leveraging ZIQI-Eval, we conduct a comprehensive evaluation over 16 LLMs to evaluate and analyze LLMs' performance in the domain of music. Results indicate that all LLMs perform poorly on the ZIQI-Eval benchmark, suggesting significant room for improvement in their musical capabilities. With ZIQI-Eval, we aim to provide a standardized and robust evaluation framework that facilitates a comprehensive assessment of LLMs' music-related abilities. The dataset is available at GitHub\footnote{https://github.com/zcli-charlie/ZIQI-Eval} and HuggingFace\footnote{https://huggingface.co/datasets/MYTH-Lab/ZIQI-Eval}.
title The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2406.15885