persoDA: Personalized Data Augmentation for Personalized ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Parada, Pablo Peso, Fontalis, Spyros, Jalal, Md Asif, Saravanan, Karthikeyan, Drosou, Anastasios, Ozay, Mete, Lee, Gil Ho, Lee, Jungin, Jung, Seokyeong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912192585007104
author Parada, Pablo Peso
Fontalis, Spyros
Jalal, Md Asif
Saravanan, Karthikeyan
Drosou, Anastasios
Ozay, Mete
Lee, Gil Ho
Lee, Jungin
Jung, Seokyeong
author_facet Parada, Pablo Peso
Fontalis, Spyros
Jalal, Md Asif
Saravanan, Karthikeyan
Drosou, Anastasios
Ozay, Mete
Lee, Gil Ho
Lee, Jungin
Jung, Seokyeong
contents Data augmentation (DA) is ubiquitously used in training of Automatic Speech Recognition (ASR) models. DA offers increased data variability, robustness and generalization against different acoustic distortions. Recently, personalization of ASR models on mobile devices has been shown to improve Word Error Rate (WER). This paper evaluates data augmentation in this context and proposes persoDA; a DA method driven by user's data utilized to personalize ASR. persoDA aims to augment training with data specifically tuned towards acoustic characteristics of the end-user, as opposed to standard augmentation based on Multi-Condition Training (MCT) that applies random reverberation and noises. Our evaluation with an ASR conformer-based baseline trained on Librispeech and personalized for VOICES shows that persoDA achieves a 13.9% relative WER reduction over using standard data augmentation (using random noise & reverberation). Furthermore, persoDA shows 16% to 20% faster convergence over MCT.
format Preprint
id arxiv_https___arxiv_org_abs_2501_09113
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle persoDA: Personalized Data Augmentation for Personalized ASR
Parada, Pablo Peso
Fontalis, Spyros
Jalal, Md Asif
Saravanan, Karthikeyan
Drosou, Anastasios
Ozay, Mete
Lee, Gil Ho
Lee, Jungin
Jung, Seokyeong
Audio and Speech Processing
Sound
Data augmentation (DA) is ubiquitously used in training of Automatic Speech Recognition (ASR) models. DA offers increased data variability, robustness and generalization against different acoustic distortions. Recently, personalization of ASR models on mobile devices has been shown to improve Word Error Rate (WER). This paper evaluates data augmentation in this context and proposes persoDA; a DA method driven by user's data utilized to personalize ASR. persoDA aims to augment training with data specifically tuned towards acoustic characteristics of the end-user, as opposed to standard augmentation based on Multi-Condition Training (MCT) that applies random reverberation and noises. Our evaluation with an ASR conformer-based baseline trained on Librispeech and personalized for VOICES shows that persoDA achieves a 13.9% relative WER reduction over using standard data augmentation (using random noise & reverberation). Furthermore, persoDA shows 16% to 20% faster convergence over MCT.
title persoDA: Personalized Data Augmentation for Personalized ASR
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2501.09113