Saved in:
Bibliographic Details
Main Authors: Polok, Alexander, Medennikov, Ivan, Černocký, Jan, Watanabe, Shinji, Burget, Lukáš, Cornell, Samuele
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.15442
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913131816550400
author Polok, Alexander
Medennikov, Ivan
Černocký, Jan
Watanabe, Shinji
Burget, Lukáš
Cornell, Samuele
author_facet Polok, Alexander
Medennikov, Ivan
Černocký, Jan
Watanabe, Shinji
Burget, Lukáš
Cornell, Samuele
contents Recent breakthroughs in multi-talker ASR (MT-ASR) and speaker diarization (SD) rely on synthetic data to mitigate the scarcity of large-scale conversational recordings, yet the impact of specific simulation choices remains poorly understood. To mind the gap between simulated mixtures and real-world interactions, we present a study of synthetic data generation for leading MT-ASR (DiCoW) and SD (Sortformer) systems. By introducing FastMSS, a highly efficient open-source simulator, we analyze turn-taking dynamics, source domain, acoustic augmentation, and data mixing strategies. Our findings reveal that optimal simulation recipes are highly task-dependent: increasing speech overlap benefits ASR but degrades diarization. Furthermore, broad source diversity consistently outperforms exact domain matching. Ultimately, synthetic-only training approaches real-data baselines, and combining simulated data with real recordings yields substantial gains over real-only training across both tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15442
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
Polok, Alexander
Medennikov, Ivan
Černocký, Jan
Watanabe, Shinji
Burget, Lukáš
Cornell, Samuele
Audio and Speech Processing
Recent breakthroughs in multi-talker ASR (MT-ASR) and speaker diarization (SD) rely on synthetic data to mitigate the scarcity of large-scale conversational recordings, yet the impact of specific simulation choices remains poorly understood. To mind the gap between simulated mixtures and real-world interactions, we present a study of synthetic data generation for leading MT-ASR (DiCoW) and SD (Sortformer) systems. By introducing FastMSS, a highly efficient open-source simulator, we analyze turn-taking dynamics, source domain, acoustic augmentation, and data mixing strategies. Our findings reveal that optimal simulation recipes are highly task-dependent: increasing speech overlap benefits ASR but degrades diarization. Furthermore, broad source diversity consistently outperforms exact domain matching. Ultimately, synthetic-only training approaches real-data baselines, and combining simulated data with real recordings yields substantial gains over real-only training across both tasks.
title Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
topic Audio and Speech Processing
url https://arxiv.org/abs/2605.15442