Improving Self-supervised Pre-training using Accent-Specific Codebooks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Prabhu, Darshan, Gupta, Abhishek, Nitsure, Omkar, Jyothi, Preethi, Ganapathy, Sriram
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909241563938816
author Prabhu, Darshan
Gupta, Abhishek
Nitsure, Omkar
Jyothi, Preethi
Ganapathy, Sriram
author_facet Prabhu, Darshan
Gupta, Abhishek
Nitsure, Omkar
Jyothi, Preethi
Ganapathy, Sriram
contents Speech accents present a serious challenge to the performance of state-of-the-art end-to-end Automatic Speech Recognition (ASR) systems. Even with self-supervised learning and pre-training of ASR models, accent invariance is seldom achieved. In this work, we propose an accent-aware adaptation technique for self-supervised learning that introduces a trainable set of accent-specific codebooks to the self-supervised architecture. These learnable codebooks enable the model to capture accent specific information during pre-training, that is further refined during ASR finetuning. On the Mozilla Common Voice dataset, our proposed approach outperforms all other accent-adaptation approaches on both seen and unseen English accents, with up to 9% relative reduction in word error rate (WER).
format Preprint
id arxiv_https___arxiv_org_abs_2407_03734
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Self-supervised Pre-training using Accent-Specific Codebooks
Prabhu, Darshan
Gupta, Abhishek
Nitsure, Omkar
Jyothi, Preethi
Ganapathy, Sriram
Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
Speech accents present a serious challenge to the performance of state-of-the-art end-to-end Automatic Speech Recognition (ASR) systems. Even with self-supervised learning and pre-training of ASR models, accent invariance is seldom achieved. In this work, we propose an accent-aware adaptation technique for self-supervised learning that introduces a trainable set of accent-specific codebooks to the self-supervised architecture. These learnable codebooks enable the model to capture accent specific information during pre-training, that is further refined during ASR finetuning. On the Mozilla Common Voice dataset, our proposed approach outperforms all other accent-adaptation approaches on both seen and unseen English accents, with up to 9% relative reduction in word error rate (WER).
title Improving Self-supervised Pre-training using Accent-Specific Codebooks
topic Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2407.03734