SingMOS-Pro: An Comprehensive Benchmark for Singing Quality Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Yuxun, Liu, Lan, Feng, Wenhao, Zhao, Yiwen, Han, Jionghao, Yu, Yifeng, Shi, Jiatong, Jin, Qin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915757235896320
author Tang, Yuxun
Liu, Lan
Feng, Wenhao
Zhao, Yiwen
Han, Jionghao
Yu, Yifeng
Shi, Jiatong
Jin, Qin
author_facet Tang, Yuxun
Liu, Lan
Feng, Wenhao
Zhao, Yiwen
Han, Jionghao
Yu, Yifeng
Shi, Jiatong
Jin, Qin
contents Singing voice generation progresses rapidly, yet evaluating singing quality remains a critical challenge. Human subjective assessment, typically in the form of listening tests, is costly and time consuming, while existing objective metrics capture only limited perceptual aspects. In this work, we introduce SingMOS-Pro, a dataset for automatic singing quality assessment. Building on our preview version SingMOS, which provides only overall ratings, SingMOS-Pro extends the annotations of the additional data to include lyrics, melody, and overall quality, offering broader coverage and greater diversity. The dataset contains 7,981 singing clips generated by 41 models across 12 datasets, spanning from early systems to recent state-of-the-art approaches. Each clip is rated by at least five experienced annotators to ensure reliability and consistency. Furthermore, we investigate strategies for effectively utilizing MOS data annotated under heterogeneous standards and benchmark several widely used evaluation methods from related tasks on SingMOS-Pro, establishing strong baselines and practical references for future research. The dataset is publicly available at https://huggingface.co/datasets/TangRain/SingMOS-Pro.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01812
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SingMOS-Pro: An Comprehensive Benchmark for Singing Quality Assessment
Tang, Yuxun
Liu, Lan
Feng, Wenhao
Zhao, Yiwen
Han, Jionghao
Yu, Yifeng
Shi, Jiatong
Jin, Qin
Sound
Artificial Intelligence
Audio and Speech Processing
Singing voice generation progresses rapidly, yet evaluating singing quality remains a critical challenge. Human subjective assessment, typically in the form of listening tests, is costly and time consuming, while existing objective metrics capture only limited perceptual aspects. In this work, we introduce SingMOS-Pro, a dataset for automatic singing quality assessment. Building on our preview version SingMOS, which provides only overall ratings, SingMOS-Pro extends the annotations of the additional data to include lyrics, melody, and overall quality, offering broader coverage and greater diversity. The dataset contains 7,981 singing clips generated by 41 models across 12 datasets, spanning from early systems to recent state-of-the-art approaches. Each clip is rated by at least five experienced annotators to ensure reliability and consistency. Furthermore, we investigate strategies for effectively utilizing MOS data annotated under heterogeneous standards and benchmark several widely used evaluation methods from related tasks on SingMOS-Pro, establishing strong baselines and practical references for future research. The dataset is publicly available at https://huggingface.co/datasets/TangRain/SingMOS-Pro.
title SingMOS-Pro: An Comprehensive Benchmark for Singing Quality Assessment
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2510.01812