SOT Triggered Neural Clustering for Speaker Attributed ASR

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zheng, Xianrui, Sun, Guangzhi, Zhang, Chao, Woodland, Philip C.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912009746907136
author Zheng, Xianrui
Sun, Guangzhi
Zhang, Chao
Woodland, Philip C.
author_facet Zheng, Xianrui
Sun, Guangzhi
Zhang, Chao
Woodland, Philip C.
contents This paper introduces a novel approach to speaker-attributed ASR transcription using a neural clustering method. With a parallel processing mechanism, diarisation and ASR can be applied simultaneously, helping to prevent the accumulation of errors from one sub-system to the next in a cascaded system. This is achieved by the use of ASR, trained using a serialised output training method, together with segment-level discriminative neural clustering (SDNC) to assign speaker labels. With SDNC, our system does not require an extra non-neural clustering method to assign speaker labels, thus allowing the entire system to be based on neural networks. Experimental results on the AMI meeting dataset demonstrate that SDNC outperforms spectral clustering (SC) by a 19% relative diarisation error rate (DER) reduction on the AMI Eval set. When compared with the cascaded system with SC, the parallel system with SDNC gives a 7%/4% relative improvement in cpWER on the Dev/Eval set.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02007
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SOT Triggered Neural Clustering for Speaker Attributed ASR
Zheng, Xianrui
Sun, Guangzhi
Zhang, Chao
Woodland, Philip C.
Audio and Speech Processing
This paper introduces a novel approach to speaker-attributed ASR transcription using a neural clustering method. With a parallel processing mechanism, diarisation and ASR can be applied simultaneously, helping to prevent the accumulation of errors from one sub-system to the next in a cascaded system. This is achieved by the use of ASR, trained using a serialised output training method, together with segment-level discriminative neural clustering (SDNC) to assign speaker labels. With SDNC, our system does not require an extra non-neural clustering method to assign speaker labels, thus allowing the entire system to be based on neural networks. Experimental results on the AMI meeting dataset demonstrate that SDNC outperforms spectral clustering (SC) by a 19% relative diarisation error rate (DER) reduction on the AMI Eval set. When compared with the cascaded system with SC, the parallel system with SDNC gives a 7%/4% relative improvement in cpWER on the Dev/Eval set.
title SOT Triggered Neural Clustering for Speaker Attributed ASR
topic Audio and Speech Processing
url https://arxiv.org/abs/2407.02007