persoDA: Personalized Data Augmentation for Personalized ASR
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912192585007104 |
|---|---|
| author | Parada, Pablo Peso Fontalis, Spyros Jalal, Md Asif Saravanan, Karthikeyan Drosou, Anastasios Ozay, Mete Lee, Gil Ho Lee, Jungin Jung, Seokyeong |
| author_facet | Parada, Pablo Peso Fontalis, Spyros Jalal, Md Asif Saravanan, Karthikeyan Drosou, Anastasios Ozay, Mete Lee, Gil Ho Lee, Jungin Jung, Seokyeong |
| contents | Data augmentation (DA) is ubiquitously used in training of Automatic Speech Recognition (ASR) models. DA offers increased data variability, robustness and generalization against different acoustic distortions. Recently, personalization of ASR models on mobile devices has been shown to improve Word Error Rate (WER). This paper evaluates data augmentation in this context and proposes persoDA; a DA method driven by user's data utilized to personalize ASR. persoDA aims to augment training with data specifically tuned towards acoustic characteristics of the end-user, as opposed to standard augmentation based on Multi-Condition Training (MCT) that applies random reverberation and noises. Our evaluation with an ASR conformer-based baseline trained on Librispeech and personalized for VOICES shows that persoDA achieves a 13.9% relative WER reduction over using standard data augmentation (using random noise & reverberation). Furthermore, persoDA shows 16% to 20% faster convergence over MCT. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_09113 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | persoDA: Personalized Data Augmentation for Personalized ASR Parada, Pablo Peso Fontalis, Spyros Jalal, Md Asif Saravanan, Karthikeyan Drosou, Anastasios Ozay, Mete Lee, Gil Ho Lee, Jungin Jung, Seokyeong Audio and Speech Processing Sound Data augmentation (DA) is ubiquitously used in training of Automatic Speech Recognition (ASR) models. DA offers increased data variability, robustness and generalization against different acoustic distortions. Recently, personalization of ASR models on mobile devices has been shown to improve Word Error Rate (WER). This paper evaluates data augmentation in this context and proposes persoDA; a DA method driven by user's data utilized to personalize ASR. persoDA aims to augment training with data specifically tuned towards acoustic characteristics of the end-user, as opposed to standard augmentation based on Multi-Condition Training (MCT) that applies random reverberation and noises. Our evaluation with an ASR conformer-based baseline trained on Librispeech and personalized for VOICES shows that persoDA achieves a 13.9% relative WER reduction over using standard data augmentation (using random noise & reverberation). Furthermore, persoDA shows 16% to 20% faster convergence over MCT. |
| title | persoDA: Personalized Data Augmentation for Personalized ASR |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2501.09113 |