On Speaker Attribution with SURT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Raj, Desh, Wiesner, Matthew, Maciejewski, Matthew, Garcia-Perera, Leibny Paola, Povey, Daniel, Khudanpur, Sanjeev
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911766345154560
author Raj, Desh
Wiesner, Matthew
Maciejewski, Matthew
Garcia-Perera, Leibny Paola
Povey, Daniel
Khudanpur, Sanjeev
author_facet Raj, Desh
Wiesner, Matthew
Maciejewski, Matthew
Garcia-Perera, Leibny Paola
Povey, Daniel
Khudanpur, Sanjeev
contents The Streaming Unmixing and Recognition Transducer (SURT) has recently become a popular framework for continuous, streaming, multi-talker speech recognition (ASR). With advances in architecture, objectives, and mixture simulation methods, it was demonstrated that SURT can be an efficient streaming method for speaker-agnostic transcription of real meetings. In this work, we push this framework further by proposing methods to perform speaker-attributed transcription with SURT, for both short mixtures and long recordings. We achieve this by adding an auxiliary speaker branch to SURT, and synchronizing its label prediction with ASR token prediction through HAT-style blank factorization. In order to ensure consistency in relative speaker labels across different utterance groups in a recording, we propose "speaker prefixing" -- appending each chunk with high-confidence frames of speakers identified in previous chunks, to establish the relative order. We perform extensive ablation experiments on synthetic LibriSpeech mixtures to validate our design choices, and demonstrate the efficacy of our final model on the AMI corpus.
format Preprint
id arxiv_https___arxiv_org_abs_2401_15676
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On Speaker Attribution with SURT
Raj, Desh
Wiesner, Matthew
Maciejewski, Matthew
Garcia-Perera, Leibny Paola
Povey, Daniel
Khudanpur, Sanjeev
Audio and Speech Processing
Sound
The Streaming Unmixing and Recognition Transducer (SURT) has recently become a popular framework for continuous, streaming, multi-talker speech recognition (ASR). With advances in architecture, objectives, and mixture simulation methods, it was demonstrated that SURT can be an efficient streaming method for speaker-agnostic transcription of real meetings. In this work, we push this framework further by proposing methods to perform speaker-attributed transcription with SURT, for both short mixtures and long recordings. We achieve this by adding an auxiliary speaker branch to SURT, and synchronizing its label prediction with ASR token prediction through HAT-style blank factorization. In order to ensure consistency in relative speaker labels across different utterance groups in a recording, we propose "speaker prefixing" -- appending each chunk with high-confidence frames of speakers identified in previous chunks, to establish the relative order. We perform extensive ablation experiments on synthetic LibriSpeech mixtures to validate our design choices, and demonstrate the efficacy of our final model on the AMI corpus.
title On Speaker Attribution with SURT
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2401.15676