DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent Adaptation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kothawade, Suraj, Mekala, Anmol, D, Chandra Sekhara, Kothyari, Mayank, Iyer, Rishabh, Ramakrishnan, Ganesh, Jyothi, Preethi
Format: Preprint
Veröffentlicht: 2021
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913614850424832
author Kothawade, Suraj
Mekala, Anmol
D, Chandra Sekhara
Kothyari, Mayank
Iyer, Rishabh
Ramakrishnan, Ganesh
Jyothi, Preethi
author_facet Kothawade, Suraj
Mekala, Anmol
D, Chandra Sekhara
Kothyari, Mayank
Iyer, Rishabh
Ramakrishnan, Ganesh
Jyothi, Preethi
contents State-of-the-art Automatic Speech Recognition (ASR) systems are known to exhibit disparate performance on varying speech accents. To improve performance on a specific target accent, a commonly adopted solution is to finetune the ASR model using accent-specific labeled speech. However, acquiring large amounts of labeled speech for specific target accents is challenging. Choosing an informative subset of speech samples that are most representative of the target accents becomes important for effective ASR finetuning. To address this problem, we propose DITTO (Data-efficient and faIr Targeted subseT selectiOn) that uses Submodular Mutual Information (SMI) functions as acquisition functions to find the most informative set of utterances matching a target accent within a fixed budget. An important feature of DITTO is that it supports fair targeting for multiple accents, i.e. it can automatically select representative data points from multiple accents when the ASR model needs to perform well on more than one accent. We show that DITTO is 3-5 times more label-efficient than other speech selection methods on the IndicTTS and L2 datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2110_04908
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent Adaptation
Kothawade, Suraj
Mekala, Anmol
D, Chandra Sekhara
Kothyari, Mayank
Iyer, Rishabh
Ramakrishnan, Ganesh
Jyothi, Preethi
Audio and Speech Processing
Sound
State-of-the-art Automatic Speech Recognition (ASR) systems are known to exhibit disparate performance on varying speech accents. To improve performance on a specific target accent, a commonly adopted solution is to finetune the ASR model using accent-specific labeled speech. However, acquiring large amounts of labeled speech for specific target accents is challenging. Choosing an informative subset of speech samples that are most representative of the target accents becomes important for effective ASR finetuning. To address this problem, we propose DITTO (Data-efficient and faIr Targeted subseT selectiOn) that uses Submodular Mutual Information (SMI) functions as acquisition functions to find the most informative set of utterances matching a target accent within a fixed budget. An important feature of DITTO is that it supports fair targeting for multiple accents, i.e. it can automatically select representative data points from multiple accents when the ASR model needs to perform well on more than one accent. We show that DITTO is 3-5 times more label-efficient than other speech selection methods on the IndicTTS and L2 datasets.
title DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent Adaptation
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2110.04908