Consistency Based Unsupervised Self-training For ASR Personalisation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917572284252160 |
|---|---|
| author | Zhang, Jisi Rajan, Vandana Mehmood, Haaris Tuckey, David Parada, Pablo Peso Jalal, Md Asif Saravanan, Karthikeyan Lee, Gil Ho Lee, Jungin Jung, Seokyeong |
| author_facet | Zhang, Jisi Rajan, Vandana Mehmood, Haaris Tuckey, David Parada, Pablo Peso Jalal, Md Asif Saravanan, Karthikeyan Lee, Gil Ho Lee, Jungin Jung, Seokyeong |
| contents | On-device Automatic Speech Recognition (ASR) models trained on speech data of a large population might underperform for individuals unseen during training. This is due to a domain shift between user data and the original training data, differed by user's speaking characteristics and environmental acoustic conditions. ASR personalisation is a solution that aims to exploit user data to improve model robustness. The majority of ASR personalisation methods assume labelled user data for supervision. Personalisation without any labelled data is challenging due to limited data size and poor quality of recorded audio samples. This work addresses unsupervised personalisation by developing a novel consistency based training method via pseudo-labelling. Our method achieves a relative Word Error Rate Reduction (WERR) of 17.3% on unlabelled training data and 8.1% on held-out data compared to a pre-trained model, and outperforms the current state-of-the art methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_12085 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Consistency Based Unsupervised Self-training For ASR Personalisation Zhang, Jisi Rajan, Vandana Mehmood, Haaris Tuckey, David Parada, Pablo Peso Jalal, Md Asif Saravanan, Karthikeyan Lee, Gil Ho Lee, Jungin Jung, Seokyeong Audio and Speech Processing Sound On-device Automatic Speech Recognition (ASR) models trained on speech data of a large population might underperform for individuals unseen during training. This is due to a domain shift between user data and the original training data, differed by user's speaking characteristics and environmental acoustic conditions. ASR personalisation is a solution that aims to exploit user data to improve model robustness. The majority of ASR personalisation methods assume labelled user data for supervision. Personalisation without any labelled data is challenging due to limited data size and poor quality of recorded audio samples. This work addresses unsupervised personalisation by developing a novel consistency based training method via pseudo-labelling. Our method achieves a relative Word Error Rate Reduction (WERR) of 17.3% on unlabelled training data and 8.1% on held-out data compared to a pre-trained model, and outperforms the current state-of-the art methods. |
| title | Consistency Based Unsupervised Self-training For ASR Personalisation |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2401.12085 |