Modeling Overlapped Speech with Shuffles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wiesner, Matthew, Cornell, Samuele, Polok, Alexander, Yang, Lucas Ondel, Burget, Lukáš, Khudanpur, Sanjeev
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914408194637824
author Wiesner, Matthew
Cornell, Samuele
Polok, Alexander
Yang, Lucas Ondel
Burget, Lukáš
Khudanpur, Sanjeev
author_facet Wiesner, Matthew
Cornell, Samuele
Polok, Alexander
Yang, Lucas Ondel
Burget, Lukáš
Khudanpur, Sanjeev
contents We propose to model parallel streams of data, such as overlapped speech, using shuffles. Specifically, this paper shows how the shuffle product and partial order finite-state automata (FSAs) can be used for alignment and speaker-attributed transcription of overlapped speech. We train using the total score on these FSAs as a loss function, marginalizing over all possible serializations of overlapping sequences at subword, word, and phrase levels. To reduce graph size, we impose temporal constraints by constructing partial order FSAs. We address speaker attribution by modeling (token, speaker) tuples directly. Viterbi alignment through the shuffle product FSA directly enables one-pass alignment. We evaluate performance on synthetic LibriSpeech overlaps. To our knowledge, this is the first algorithm that enables single-pass alignment of multi-talker recordings. All algorithms are implemented using k2 / Icefall.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17769
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Modeling Overlapped Speech with Shuffles
Wiesner, Matthew
Cornell, Samuele
Polok, Alexander
Yang, Lucas Ondel
Burget, Lukáš
Khudanpur, Sanjeev
Sound
Computation and Language
Machine Learning
Audio and Speech Processing
We propose to model parallel streams of data, such as overlapped speech, using shuffles. Specifically, this paper shows how the shuffle product and partial order finite-state automata (FSAs) can be used for alignment and speaker-attributed transcription of overlapped speech. We train using the total score on these FSAs as a loss function, marginalizing over all possible serializations of overlapping sequences at subword, word, and phrase levels. To reduce graph size, we impose temporal constraints by constructing partial order FSAs. We address speaker attribution by modeling (token, speaker) tuples directly. Viterbi alignment through the shuffle product FSA directly enables one-pass alignment. We evaluate performance on synthetic LibriSpeech overlaps. To our knowledge, this is the first algorithm that enables single-pass alignment of multi-talker recordings. All algorithms are implemented using k2 / Icefall.
title Modeling Overlapped Speech with Shuffles
topic Sound
Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2603.17769