Enhanced Reverberation as Supervision for Unsupervised Speech Separation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Saijo, Kohei, Wichern, Gordon, Germain, François G., Pan, Zexu, Roux, Jonathan Le
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911980449693696
author Saijo, Kohei
Wichern, Gordon
Germain, François G.
Pan, Zexu
Roux, Jonathan Le
author_facet Saijo, Kohei
Wichern, Gordon
Germain, François G.
Pan, Zexu
Roux, Jonathan Le
contents Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models are trained so that sources predicted from a mixture at an input channel can be mapped to reconstruct a mixture at a target channel. However, stable unsupervised training has so far only been achieved in over-determined source-channel conditions, leaving the key determined case unsolved. This work proposes enhanced RAS (ERAS) for solving this problem. Through qualitative analysis, we found that stable training can be achieved by leveraging the loss term to alleviate the frequency-permutation problem. Separation performance is also boosted by adding a novel loss term where separated signals mapped back to their own input mixture are used as pseudo-targets for the signals separated from other channels and mapped to the same channel. Experimental results demonstrate high stability and performance of ERAS.
format Preprint
id arxiv_https___arxiv_org_abs_2408_03438
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhanced Reverberation as Supervision for Unsupervised Speech Separation
Saijo, Kohei
Wichern, Gordon
Germain, François G.
Pan, Zexu
Roux, Jonathan Le
Audio and Speech Processing
Sound
Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models are trained so that sources predicted from a mixture at an input channel can be mapped to reconstruct a mixture at a target channel. However, stable unsupervised training has so far only been achieved in over-determined source-channel conditions, leaving the key determined case unsolved. This work proposes enhanced RAS (ERAS) for solving this problem. Through qualitative analysis, we found that stable training can be achieved by leveraging the loss term to alleviate the frequency-permutation problem. Separation performance is also boosted by adding a novel loss term where separated signals mapped back to their own input mixture are used as pseudo-targets for the signals separated from other channels and mapped to the same channel. Experimental results demonstrate high stability and performance of ERAS.
title Enhanced Reverberation as Supervision for Unsupervised Speech Separation
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2408.03438