Error Analysis in a Modular Meeting Transcription System

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vieting, Peter, Berger, Simon, von Neumann, Thilo, Boeddeker, Christoph, Schlüter, Ralf, Haeb-Umbach, Reinhold
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909783742742528
author Vieting, Peter
Berger, Simon
von Neumann, Thilo
Boeddeker, Christoph
Schlüter, Ralf
Haeb-Umbach, Reinhold
author_facet Vieting, Peter
Berger, Simon
von Neumann, Thilo
Boeddeker, Christoph
Schlüter, Ralf
Haeb-Umbach, Reinhold
contents Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previously proposed framework for analyzing leakage in speech separation with proper sensitivity to temporal locality. We show that there is significant leakage to the cross channel in areas where only the primary speaker is active. At the same time, the results demonstrate that this does not affect the final performance much as these leaked parts are largely ignored by the voice activity detection (VAD). Furthermore, different segmentations are compared showing that advanced diarization approaches are able to reduce the gap to oracle segmentation by a third compared to a simple energy-based VAD. We additionally reveal what factors contribute to the remaining difference. The results represent state-of-the-art performance on LibriCSS among systems that train the recognition module on LibriSpeech data only.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10143
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Error Analysis in a Modular Meeting Transcription System
Vieting, Peter
Berger, Simon
von Neumann, Thilo
Boeddeker, Christoph
Schlüter, Ralf
Haeb-Umbach, Reinhold
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previously proposed framework for analyzing leakage in speech separation with proper sensitivity to temporal locality. We show that there is significant leakage to the cross channel in areas where only the primary speaker is active. At the same time, the results demonstrate that this does not affect the final performance much as these leaked parts are largely ignored by the voice activity detection (VAD). Furthermore, different segmentations are compared showing that advanced diarization approaches are able to reduce the gap to oracle segmentation by a third compared to a simple energy-based VAD. We additionally reveal what factors contribute to the remaining difference. The results represent state-of-the-art performance on LibriCSS among systems that train the recognition module on LibriSpeech data only.
title Error Analysis in a Modular Meeting Transcription System
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2509.10143