Using RLHF to align speech enhancement approaches to mean-opinion quality scores

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kumar, Anurag, Perrault, Andrew, Williamson, Donald S.
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909353057976320
author Kumar, Anurag
Perrault, Andrew
Williamson, Donald S.
author_facet Kumar, Anurag
Perrault, Andrew
Williamson, Donald S.
contents Objective speech quality measures are typically used to assess speech enhancement algorithms, but it has been shown that they are sub-optimal as learning objectives because they do not always align well with human subjective ratings. This misalignment often results in noticeable distortions and artifacts that cause speech enhancement to be ineffective. To address these issues, we propose a reinforcement learning from human feedback (RLHF) framework to fine-tune an existing speech enhancement approach by optimizing performance using a mean-opinion score (MOS)-based reward model. Our results show that the RLHF-finetuned model has the best performance across different benchmarks for both objective and MOS-based speech quality assessment metrics on the Voicebank+DEMAND dataset. Through ablation studies, we show that both policy gradient loss and supervised MSE loss are important for balanced optimization across the different metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13182
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Using RLHF to align speech enhancement approaches to mean-opinion quality scores
Kumar, Anurag
Perrault, Andrew
Williamson, Donald S.
Audio and Speech Processing
Sound
Objective speech quality measures are typically used to assess speech enhancement algorithms, but it has been shown that they are sub-optimal as learning objectives because they do not always align well with human subjective ratings. This misalignment often results in noticeable distortions and artifacts that cause speech enhancement to be ineffective. To address these issues, we propose a reinforcement learning from human feedback (RLHF) framework to fine-tune an existing speech enhancement approach by optimizing performance using a mean-opinion score (MOS)-based reward model. Our results show that the RLHF-finetuned model has the best performance across different benchmarks for both objective and MOS-based speech quality assessment metrics on the Voicebank+DEMAND dataset. Through ablation studies, we show that both policy gradient loss and supervised MSE loss are important for balanced optimization across the different metrics.
title Using RLHF to align speech enhancement approaches to mean-opinion quality scores
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2410.13182