Joint ASR and Speaker Role Tagging with Serialized Output Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Anfeng, Feng, Tiantian, Narayanan, Shrikanth
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915338892869632
author Xu, Anfeng
Feng, Tiantian
Narayanan, Shrikanth
author_facet Xu, Anfeng
Feng, Tiantian
Narayanan, Shrikanth
contents Automatic Speech Recognition systems have made significant progress with large-scale pre-trained models. However, most current systems focus solely on transcribing the speech without identifying speaker roles, a function that is critical for conversational AI. In this work, we investigate the use of serialized output training (SOT) for joint ASR and speaker role tagging. By augmenting Whisper with role-specific tokens and fine-tuning it with SOT, we enable the model to generate role-aware transcriptions in a single decoding pass. We compare the SOT approach against a self-supervised previous baseline method on two real-world conversational datasets. Our findings show that this approach achieves more than 10% reduction in multi-talker WER, demonstrating its feasibility as a unified model for speaker-role aware speech transcription.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10349
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Joint ASR and Speaker Role Tagging with Serialized Output Training
Xu, Anfeng
Feng, Tiantian
Narayanan, Shrikanth
Audio and Speech Processing
Sound
Automatic Speech Recognition systems have made significant progress with large-scale pre-trained models. However, most current systems focus solely on transcribing the speech without identifying speaker roles, a function that is critical for conversational AI. In this work, we investigate the use of serialized output training (SOT) for joint ASR and speaker role tagging. By augmenting Whisper with role-specific tokens and fine-tuning it with SOT, we enable the model to generate role-aware transcriptions in a single decoding pass. We compare the SOT approach against a self-supervised previous baseline method on two real-world conversational datasets. Our findings show that this approach achieves more than 10% reduction in multi-talker WER, demonstrating its feasibility as a unified model for speaker-role aware speech transcription.
title Joint ASR and Speaker Role Tagging with Serialized Output Training
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.10349