Improving Self-supervised Pre-training using Accent-Specific Codebooks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866909241563938816 |
|---|---|
| author | Prabhu, Darshan Gupta, Abhishek Nitsure, Omkar Jyothi, Preethi Ganapathy, Sriram |
| author_facet | Prabhu, Darshan Gupta, Abhishek Nitsure, Omkar Jyothi, Preethi Ganapathy, Sriram |
| contents | Speech accents present a serious challenge to the performance of state-of-the-art end-to-end Automatic Speech Recognition (ASR) systems. Even with self-supervised learning and pre-training of ASR models, accent invariance is seldom achieved. In this work, we propose an accent-aware adaptation technique for self-supervised learning that introduces a trainable set of accent-specific codebooks to the self-supervised architecture. These learnable codebooks enable the model to capture accent specific information during pre-training, that is further refined during ASR finetuning. On the Mozilla Common Voice dataset, our proposed approach outperforms all other accent-adaptation approaches on both seen and unseen English accents, with up to 9% relative reduction in word error rate (WER). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_03734 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Improving Self-supervised Pre-training using Accent-Specific Codebooks Prabhu, Darshan Gupta, Abhishek Nitsure, Omkar Jyothi, Preethi Ganapathy, Sriram Computation and Language Artificial Intelligence Machine Learning Sound Audio and Speech Processing Speech accents present a serious challenge to the performance of state-of-the-art end-to-end Automatic Speech Recognition (ASR) systems. Even with self-supervised learning and pre-training of ASR models, accent invariance is seldom achieved. In this work, we propose an accent-aware adaptation technique for self-supervised learning that introduces a trainable set of accent-specific codebooks to the self-supervised architecture. These learnable codebooks enable the model to capture accent specific information during pre-training, that is further refined during ASR finetuning. On the Mozilla Common Voice dataset, our proposed approach outperforms all other accent-adaptation approaches on both seen and unseen English accents, with up to 9% relative reduction in word error rate (WER). |
| title | Improving Self-supervised Pre-training using Accent-Specific Codebooks |
| topic | Computation and Language Artificial Intelligence Machine Learning Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2407.03734 |