Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910846028873728 |
|---|---|
| author | Vieting, Peter Berger, Simon von Neumann, Thilo Boeddeker, Christoph Schlüter, Ralf Haeb-Umbach, Reinhold |
| author_facet | Vieting, Peter Berger, Simon von Neumann, Thilo Boeddeker, Christoph Schlüter, Ralf Haeb-Umbach, Reinhold |
| contents | Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recently, TF-GridNet has shown impressive performance in speech separation in real reverberant conditions. Furthermore, a mixture encoder was proposed that leverages the mixed speech to mitigate the effect of separation artifacts. In this work, we extended the mixture encoder from a static two-speaker scenario to a natural meeting context featuring an arbitrary number of speakers and varying degrees of overlap. We further demonstrate its limits by the integration with separators of varying strength including TF-GridNet. Our experiments result in a new state-of-the-art performance on LibriCSS using a single microphone. They show that TF-GridNet largely closes the gap between previous methods and oracle separation independent of mixture encoding. We further investigate the remaining potential for improvement. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2309_08454 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription Vieting, Peter Berger, Simon von Neumann, Thilo Boeddeker, Christoph Schlüter, Ralf Haeb-Umbach, Reinhold Audio and Speech Processing Computation and Language Machine Learning Sound Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recently, TF-GridNet has shown impressive performance in speech separation in real reverberant conditions. Furthermore, a mixture encoder was proposed that leverages the mixed speech to mitigate the effect of separation artifacts. In this work, we extended the mixture encoder from a static two-speaker scenario to a natural meeting context featuring an arbitrary number of speakers and varying degrees of overlap. We further demonstrate its limits by the integration with separators of varying strength including TF-GridNet. Our experiments result in a new state-of-the-art performance on LibriCSS using a single microphone. They show that TF-GridNet largely closes the gap between previous methods and oracle separation independent of mixture encoding. We further investigate the remaining potential for improvement. |
| title | Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription |
| topic | Audio and Speech Processing Computation and Language Machine Learning Sound |
| url | https://arxiv.org/abs/2309.08454 |