Saved in:
Bibliographic Details
Main Authors: Ullah, Asad, Ragano, Alessandro, Hines, Andrew
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2309.12763
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914851396255744
author Ullah, Asad
Ragano, Alessandro
Hines, Andrew
author_facet Ullah, Asad
Ragano, Alessandro
Hines, Andrew
contents Self-supervised representation learning (SSRL) has demonstrated superior performance than supervised models for tasks including phoneme recognition. Training SSRL models poses a challenge for low-resource languages where sufficient pre-training data may not be available. A common approach is cross-lingual pre-training. Instead, we propose to use audio augmentation techniques, namely: pitch variation, noise addition, accented target language and other language speech to pre-train SSRL models in a low resource condition and evaluate phoneme recognition. Our comparisons found that a combined synthetic augmentations (noise/pitch) strategy outperformed accent and language knowledge transfer. Furthermore, we examined the scaling factor of augmented data to achieve equivalent performance to model pre-trained with target domain speech. Our findings suggest that for resource-constrained languages, combined augmentations can be a viable option than other augmentations.
format Preprint
id arxiv_https___arxiv_org_abs_2309_12763
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Reduce, Reuse, Recycle: Is Perturbed Data better than Other Language augmentation for Low Resource Self-Supervised Speech Models
Ullah, Asad
Ragano, Alessandro
Hines, Andrew
Audio and Speech Processing
Computation and Language
Sound
Self-supervised representation learning (SSRL) has demonstrated superior performance than supervised models for tasks including phoneme recognition. Training SSRL models poses a challenge for low-resource languages where sufficient pre-training data may not be available. A common approach is cross-lingual pre-training. Instead, we propose to use audio augmentation techniques, namely: pitch variation, noise addition, accented target language and other language speech to pre-train SSRL models in a low resource condition and evaluate phoneme recognition. Our comparisons found that a combined synthetic augmentations (noise/pitch) strategy outperformed accent and language knowledge transfer. Furthermore, we examined the scaling factor of augmented data to achieve equivalent performance to model pre-trained with target domain speech. Our findings suggest that for resource-constrained languages, combined augmentations can be a viable option than other augmentations.
title Reduce, Reuse, Recycle: Is Perturbed Data better than Other Language augmentation for Low Resource Self-Supervised Speech Models
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2309.12763