Generic Speech Enhancement with Self-Supervised Representation Space Loss

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sato, Hiroshi, Ochiai, Tsubasa, Delcroix, Marc, Moriya, Takafumi, Ashihara, Takanori, Masumura, Ryo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911049283796992
author Sato, Hiroshi
Ochiai, Tsubasa
Delcroix, Marc
Moriya, Takafumi
Ashihara, Takanori
Masumura, Ryo
author_facet Sato, Hiroshi
Ochiai, Tsubasa
Delcroix, Marc
Moriya, Takafumi
Ashihara, Takanori
Masumura, Ryo
contents Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, the speech enhancement has needed to be tuned for each task. Thus, generalizing speech enhancement models to unknown downstream tasks has been challenging. This study aims to construct a generic speech enhancement front-end that can improve the performance of back-ends to solve multiple downstream tasks. To this end, we propose a novel training criterion that minimizes the distance between the enhanced and the ground truth clean signal in the feature representation domain of self-supervised learning models. Since self-supervised learning feature representations effectively express high-level speech information useful for solving various downstream tasks, the proposal is expected to make speech enhancement models preserve such information. Experimental validation demonstrates that the proposal improves the performance of multiple speech tasks while maintaining the perceptual quality of the enhanced signal.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07631
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generic Speech Enhancement with Self-Supervised Representation Space Loss
Sato, Hiroshi
Ochiai, Tsubasa
Delcroix, Marc
Moriya, Takafumi
Ashihara, Takanori
Masumura, Ryo
Audio and Speech Processing
Sound
Signal Processing
Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, the speech enhancement has needed to be tuned for each task. Thus, generalizing speech enhancement models to unknown downstream tasks has been challenging. This study aims to construct a generic speech enhancement front-end that can improve the performance of back-ends to solve multiple downstream tasks. To this end, we propose a novel training criterion that minimizes the distance between the enhanced and the ground truth clean signal in the feature representation domain of self-supervised learning models. Since self-supervised learning feature representations effectively express high-level speech information useful for solving various downstream tasks, the proposal is expected to make speech enhancement models preserve such information. Experimental validation demonstrates that the proposal improves the performance of multiple speech tasks while maintaining the perceptual quality of the enhanced signal.
title Generic Speech Enhancement with Self-Supervised Representation Space Loss
topic Audio and Speech Processing
Sound
Signal Processing
url https://arxiv.org/abs/2507.07631