Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Xiluo, Polok, Alexander, Villalba, Jesús, Thebaud, Thomas, Maciejewski, Matthew
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911191764303872
author He, Xiluo
Polok, Alexander
Villalba, Jesús
Thebaud, Thomas
Maciejewski, Matthew
author_facet He, Xiluo
Polok, Alexander
Villalba, Jesús
Thebaud, Thomas
Maciejewski, Matthew
contents An increasingly common training paradigm for multi-talker automatic speech recognition (ASR) is to use speaker activity signals to adapt single-speaker ASR models for overlapping speech. Although effective, these systems require running the ASR model once per speaker, resulting in inference costs that scale with the number of speakers and limiting their practicality. In this work, we propose a method that decouples the inference cost of activity-conditioned ASR systems from the number of speakers by converting speaker-specific activity outputs into two speaker-agnostic streams. A central challenge is that naïvely merging speaker activities into streams significantly degrades recognition, since pretrained ASR models assume contiguous, single-speaker inputs. To address this, we design new heuristics aimed at preserving conversational continuity and maintaining compatibility with existing systems. We show that our approach is compatible with Diarization-Conditioned Whisper (DiCoW) to greatly reduce runtimes on the AMI and ICSI meeting datasets while retaining competitive performance.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03630
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
He, Xiluo
Polok, Alexander
Villalba, Jesús
Thebaud, Thomas
Maciejewski, Matthew
Audio and Speech Processing
Sound
An increasingly common training paradigm for multi-talker automatic speech recognition (ASR) is to use speaker activity signals to adapt single-speaker ASR models for overlapping speech. Although effective, these systems require running the ASR model once per speaker, resulting in inference costs that scale with the number of speakers and limiting their practicality. In this work, we propose a method that decouples the inference cost of activity-conditioned ASR systems from the number of speakers by converting speaker-specific activity outputs into two speaker-agnostic streams. A central challenge is that naïvely merging speaker activities into streams significantly degrades recognition, since pretrained ASR models assume contiguous, single-speaker inputs. To address this, we design new heuristics aimed at preserving conversational continuity and maintaining compatibility with existing systems. We show that our approach is compatible with Diarization-Conditioned Whisper (DiCoW) to greatly reduce runtimes on the AMI and ICSI meeting datasets while retaining competitive performance.
title Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2510.03630