Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vieting, Peter, Berger, Simon, von Neumann, Thilo, Boeddeker, Christoph, Schlüter, Ralf, Haeb-Umbach, Reinhold
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910846028873728
author Vieting, Peter
Berger, Simon
von Neumann, Thilo
Boeddeker, Christoph
Schlüter, Ralf
Haeb-Umbach, Reinhold
author_facet Vieting, Peter
Berger, Simon
von Neumann, Thilo
Boeddeker, Christoph
Schlüter, Ralf
Haeb-Umbach, Reinhold
contents Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recently, TF-GridNet has shown impressive performance in speech separation in real reverberant conditions. Furthermore, a mixture encoder was proposed that leverages the mixed speech to mitigate the effect of separation artifacts. In this work, we extended the mixture encoder from a static two-speaker scenario to a natural meeting context featuring an arbitrary number of speakers and varying degrees of overlap. We further demonstrate its limits by the integration with separators of varying strength including TF-GridNet. Our experiments result in a new state-of-the-art performance on LibriCSS using a single microphone. They show that TF-GridNet largely closes the gap between previous methods and oracle separation independent of mixture encoding. We further investigate the remaining potential for improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2309_08454
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
Vieting, Peter
Berger, Simon
von Neumann, Thilo
Boeddeker, Christoph
Schlüter, Ralf
Haeb-Umbach, Reinhold
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recently, TF-GridNet has shown impressive performance in speech separation in real reverberant conditions. Furthermore, a mixture encoder was proposed that leverages the mixed speech to mitigate the effect of separation artifacts. In this work, we extended the mixture encoder from a static two-speaker scenario to a natural meeting context featuring an arbitrary number of speakers and varying degrees of overlap. We further demonstrate its limits by the integration with separators of varying strength including TF-GridNet. Our experiments result in a new state-of-the-art performance on LibriCSS using a single microphone. They show that TF-GridNet largely closes the gap between previous methods and oracle separation independent of mixture encoding. We further investigate the remaining potential for improvement.
title Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2309.08454