Benchmarking Music Generation Models and Metrics via Human Preference Studies

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Grötschla, Florian, Solak, Ahmet, Lanzendörfer, Luca A., Wattenhofer, Roger
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913909629255680
author Grötschla, Florian
Solak, Ahmet
Lanzendörfer, Luca A.
Wattenhofer, Roger
author_facet Grötschla, Florian
Solak, Ahmet
Lanzendörfer, Luca A.
Wattenhofer, Roger
contents Recent advancements have brought generated music closer to human-created compositions, yet evaluating these models remains challenging. While human preference is the gold standard for assessing quality, translating these subjective judgments into objective metrics, particularly for text-audio alignment and music quality, has proven difficult. In this work, we generate 6k songs using 12 state-of-the-art models and conduct a survey of 15k pairwise audio comparisons with 2.5k human participants to evaluate the correlation between human preferences and widely used metrics. To the best of our knowledge, this work is the first to rank current state-of-the-art music generation models and metrics based on human preference. To further the field of subjective metric evaluation, we provide open access to our dataset of generated music and human evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19085
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Music Generation Models and Metrics via Human Preference Studies
Grötschla, Florian
Solak, Ahmet
Lanzendörfer, Luca A.
Wattenhofer, Roger
Machine Learning
Sound
Recent advancements have brought generated music closer to human-created compositions, yet evaluating these models remains challenging. While human preference is the gold standard for assessing quality, translating these subjective judgments into objective metrics, particularly for text-audio alignment and music quality, has proven difficult. In this work, we generate 6k songs using 12 state-of-the-art models and conduct a survey of 15k pairwise audio comparisons with 2.5k human participants to evaluate the correlation between human preferences and widely used metrics. To the best of our knowledge, this work is the first to rank current state-of-the-art music generation models and metrics based on human preference. To further the field of subjective metric evaluation, we provide open access to our dataset of generated music and human evaluations.
title Benchmarking Music Generation Models and Metrics via Human Preference Studies
topic Machine Learning
Sound
url https://arxiv.org/abs/2506.19085