MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Wen-Chin, Cooper, Erica, Toda, Tomoki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910161022484480
author Huang, Wen-Chin
Cooper, Erica
Toda, Tomoki
author_facet Huang, Wen-Chin
Cooper, Erica
Toda, Tomoki
contents In this paper, we study the task of subjective speech quality assessment (SSQA), which refers to predicting the perceptual quality of speech. Owing to the development of deep neural network models, SSQA has greatly advanced and has been widely applied in scientific papers to evaluate speech generation systems. Nonetheless, the insufficient out-of-domain (OOD) generalization ability of current SSQA models is underexplored and often overlooked by researchers. To study this problem systematically, we present MOS-Bench, a diverse SSQA dataset collection that currently contains 8 training sets and 17 test sets. Through extensive experiments, we first highlight the OOD generalization challenges of existing models. We then evaluate the efficacy of multiple-dataset training, comparing straightforward data pooling against AlignNet, an existing domain-aware method. We demonstrate that pooling multiple training sets provides a simple yet effective solution, and variation in the data is a key factor for robust generalization beyond training data size.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03715
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
Huang, Wen-Chin
Cooper, Erica
Toda, Tomoki
Sound
Audio and Speech Processing
In this paper, we study the task of subjective speech quality assessment (SSQA), which refers to predicting the perceptual quality of speech. Owing to the development of deep neural network models, SSQA has greatly advanced and has been widely applied in scientific papers to evaluate speech generation systems. Nonetheless, the insufficient out-of-domain (OOD) generalization ability of current SSQA models is underexplored and often overlooked by researchers. To study this problem systematically, we present MOS-Bench, a diverse SSQA dataset collection that currently contains 8 training sets and 17 test sets. Through extensive experiments, we first highlight the OOD generalization challenges of existing models. We then evaluate the efficacy of multiple-dataset training, comparing straightforward data pooling against AlignNet, an existing domain-aware method. We demonstrate that pooling multiple training sets provides a simple yet effective solution, and variation in the data is a key factor for robust generalization beyond training data size.
title MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2411.03715