Investigating Training Objectives for Generative Speech Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Richter, Julius, de Oliveira, Danilo, Gerkmann, Timo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912192906919936
author Richter, Julius
de Oliveira, Danilo
Gerkmann, Timo
author_facet Richter, Julius
de Oliveira, Danilo
Gerkmann, Timo
contents Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This paper aims to explain the differences between these frameworks by focusing our investigation on score-based generative models and the Schrödinger bridge. We conduct a series of comprehensive experiments to compare their performance and highlight differing training behaviors. Furthermore, we propose a novel perceptual loss function tailored for the Schrödinger bridge framework, demonstrating enhanced performance and improved perceptual quality of the enhanced speech signals. All experimental code and pre-trained models are publicly available to facilitate further research and development in this domain.
format Preprint
id arxiv_https___arxiv_org_abs_2409_10753
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Investigating Training Objectives for Generative Speech Enhancement
Richter, Julius
de Oliveira, Danilo
Gerkmann, Timo
Audio and Speech Processing
Sound
Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This paper aims to explain the differences between these frameworks by focusing our investigation on score-based generative models and the Schrödinger bridge. We conduct a series of comprehensive experiments to compare their performance and highlight differing training behaviors. Furthermore, we propose a novel perceptual loss function tailored for the Schrödinger bridge framework, demonstrating enhanced performance and improved perceptual quality of the enhanced speech signals. All experimental code and pre-trained models are publicly available to facilitate further research and development in this domain.
title Investigating Training Objectives for Generative Speech Enhancement
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2409.10753