USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Guanlong, Wang, Yongqiang, Pelecanos, Jason, Zhang, Yu, Liao, Hank, Huang, Yiling, Lu, Han, Wang, Quan
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909063261978624
author Zhao, Guanlong
Wang, Yongqiang
Pelecanos, Jason
Zhang, Yu
Liao, Hank
Huang, Yiling
Lu, Han
Wang, Quan
author_facet Zhao, Guanlong
Wang, Yongqiang
Pelecanos, Jason
Zhang, Yu
Liao, Hank
Huang, Yiling
Lu, Han
Wang, Quan
contents We introduce a multilingual speaker change detection model (USM-SCD) that can simultaneously detect speaker turns and perform ASR for 96 languages. This model is adapted from a speech foundation model trained on a large quantity of supervised and unsupervised data, demonstrating the utility of fine-tuning from a large generic foundation model for a downstream task. We analyze the performance of this multilingual speaker change detection model through a series of ablation studies. We show that the USM-SCD model can achieve more than 75% average speaker change detection F1 score across a test set that consists of data from 96 languages. On American English, the USM-SCD model can achieve an 85.8% speaker change detection F1 score across various public and internal test sets, beating the previous monolingual baseline model by 21% relative. We also show that we only need to fine-tune one-quarter of the trainable model parameters to achieve the best model performance. The USM-SCD model exhibits state-of-the-art ASR quality compared with a strong public ASR baseline, making it suitable to handle both tasks with negligible additional computational cost.
format Preprint
id arxiv_https___arxiv_org_abs_2309_08023
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models
Zhao, Guanlong
Wang, Yongqiang
Pelecanos, Jason
Zhang, Yu
Liao, Hank
Huang, Yiling
Lu, Han
Wang, Quan
Audio and Speech Processing
Machine Learning
Sound
We introduce a multilingual speaker change detection model (USM-SCD) that can simultaneously detect speaker turns and perform ASR for 96 languages. This model is adapted from a speech foundation model trained on a large quantity of supervised and unsupervised data, demonstrating the utility of fine-tuning from a large generic foundation model for a downstream task. We analyze the performance of this multilingual speaker change detection model through a series of ablation studies. We show that the USM-SCD model can achieve more than 75% average speaker change detection F1 score across a test set that consists of data from 96 languages. On American English, the USM-SCD model can achieve an 85.8% speaker change detection F1 score across various public and internal test sets, beating the previous monolingual baseline model by 21% relative. We also show that we only need to fine-tune one-quarter of the trainable model parameters to achieve the best model performance. The USM-SCD model exhibits state-of-the-art ASR quality compared with a strong public ASR baseline, making it suitable to handle both tasks with negligible additional computational cost.
title USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2309.08023