Learning When to Trust Which Teacher for Weakly Supervised ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agrawal, Aakriti, Rao, Milind, Sahu, Anit Kumar, Chennupati, Gopinath, Stolcke, Andreas
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914653269917696
author Agrawal, Aakriti
Rao, Milind
Sahu, Anit Kumar
Chennupati, Gopinath
Stolcke, Andreas
author_facet Agrawal, Aakriti
Rao, Milind
Sahu, Anit Kumar
Chennupati, Gopinath
Stolcke, Andreas
contents Automatic speech recognition (ASR) training can utilize multiple experts as teacher models, each trained on a specific domain or accent. Teacher models may be opaque in nature since their architecture may be not be known or their training cadence is different from that of the student ASR model. Still, the student models are updated incrementally using the pseudo-labels generated independently by the expert teachers. In this paper, we exploit supervision from multiple domain experts in training student ASR models. This training strategy is especially useful in scenarios where few or no human transcriptions are available. To that end, we propose a Smart-Weighter mechanism that selects an appropriate expert based on the input audio, and then trains the student model in an unsupervised setting. We show the efficacy of our approach using LibriSpeech and LibriLight benchmarks and find an improvement of 4 to 25\% over baselines that uniformly weight all the experts, use a single expert model, or combine experts using ROVER.
format Preprint
id arxiv_https___arxiv_org_abs_2306_12012
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning When to Trust Which Teacher for Weakly Supervised ASR
Agrawal, Aakriti
Rao, Milind
Sahu, Anit Kumar
Chennupati, Gopinath
Stolcke, Andreas
Audio and Speech Processing
Sound
Automatic speech recognition (ASR) training can utilize multiple experts as teacher models, each trained on a specific domain or accent. Teacher models may be opaque in nature since their architecture may be not be known or their training cadence is different from that of the student ASR model. Still, the student models are updated incrementally using the pseudo-labels generated independently by the expert teachers. In this paper, we exploit supervision from multiple domain experts in training student ASR models. This training strategy is especially useful in scenarios where few or no human transcriptions are available. To that end, we propose a Smart-Weighter mechanism that selects an appropriate expert based on the input audio, and then trains the student model in an unsupervised setting. We show the efficacy of our approach using LibriSpeech and LibriLight benchmarks and find an improvement of 4 to 25\% over baselines that uniformly weight all the experts, use a single expert model, or combine experts using ROVER.
title Learning When to Trust Which Teacher for Weakly Supervised ASR
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2306.12012