ASR-Synchronized Speaker-Role Diarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghosh, Arindam, Fuhs, Mark, Kim, Bongjun, Chowdhury, Anurag, Woszczyna, Monika
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911328684212224
author Ghosh, Arindam
Fuhs, Mark
Kim, Bongjun
Chowdhury, Anurag
Woszczyna, Monika
author_facet Ghosh, Arindam
Fuhs, Mark
Kim, Bongjun
Chowdhury, Anurag
Woszczyna, Monika
contents Speaker-role diarization (RD), such as doctor vs. patient or lawyer vs. client, is practically often more useful than conventional speaker diarization (SD), which assigns only generic labels (speaker-1, speaker-2). The state-of-the-art end-to-end ASR+RD approach uses a single transducer that serializes word and role predictions (role at the end of a speaker's turn), but at the cost of degraded ASR performance. To address this, we adapt a recent joint ASR+SD framework to ASR+RD by freezing the ASR transducer and training an auxiliary RD transducer in parallel to assign a role to each ASR-predicted word. For this, we first show that SD and RD are fundamentally different tasks, exhibiting different dependencies on acoustic and linguistic information. Motivated by this, we propose (1) task-specific predictor networks and (2) using higher-layer ASR encoder features as input to the RD encoder. Additionally, we replace the blank-shared RNNT loss by cross-entropy loss along the 1-best forced-alignment path to further improve performance while reducing computational and memory requirements during RD training. Experiments on a public and a private dataset of doctor-patient conversations demonstrate that our method outperforms the best baseline with relative reductions of 6.2% and 4.5% in role-based word diarization error rate (R-WDER), respectively
format Preprint
id arxiv_https___arxiv_org_abs_2507_17765
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ASR-Synchronized Speaker-Role Diarization
Ghosh, Arindam
Fuhs, Mark
Kim, Bongjun
Chowdhury, Anurag
Woszczyna, Monika
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Speaker-role diarization (RD), such as doctor vs. patient or lawyer vs. client, is practically often more useful than conventional speaker diarization (SD), which assigns only generic labels (speaker-1, speaker-2). The state-of-the-art end-to-end ASR+RD approach uses a single transducer that serializes word and role predictions (role at the end of a speaker's turn), but at the cost of degraded ASR performance. To address this, we adapt a recent joint ASR+SD framework to ASR+RD by freezing the ASR transducer and training an auxiliary RD transducer in parallel to assign a role to each ASR-predicted word. For this, we first show that SD and RD are fundamentally different tasks, exhibiting different dependencies on acoustic and linguistic information. Motivated by this, we propose (1) task-specific predictor networks and (2) using higher-layer ASR encoder features as input to the RD encoder. Additionally, we replace the blank-shared RNNT loss by cross-entropy loss along the 1-best forced-alignment path to further improve performance while reducing computational and memory requirements during RD training. Experiments on a public and a private dataset of doctor-patient conversations demonstrate that our method outperforms the best baseline with relative reductions of 6.2% and 4.5% in role-based word diarization error rate (R-WDER), respectively
title ASR-Synchronized Speaker-Role Diarization
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.17765