Low-Resource Self-Supervised Learning with SSL-Enhanced TTS

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsu, Po-chun, Elkahky, Ali, Hsu, Wei-Ning, Adi, Yossi, Nguyen, Tu Anh, Copet, Jade, Dupoux, Emmanuel, Lee, Hung-yi, Mohamed, Abdelrahman
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917683991150592
author Hsu, Po-chun
Elkahky, Ali
Hsu, Wei-Ning
Adi, Yossi
Nguyen, Tu Anh
Copet, Jade
Dupoux, Emmanuel
Lee, Hung-yi
Mohamed, Abdelrahman
author_facet Hsu, Po-chun
Elkahky, Ali
Hsu, Wei-Ning
Adi, Yossi
Nguyen, Tu Anh
Copet, Jade
Dupoux, Emmanuel
Lee, Hung-yi
Mohamed, Abdelrahman
contents Self-supervised learning (SSL) techniques have achieved remarkable results in various speech processing tasks. Nonetheless, a significant challenge remains in reducing the reliance on vast amounts of speech data for pre-training. This paper proposes to address this challenge by leveraging synthetic speech to augment a low-resource pre-training corpus. We construct a high-quality text-to-speech (TTS) system with limited resources using SSL features and generate a large synthetic corpus for pre-training. Experimental results demonstrate that our proposed approach effectively reduces the demand for speech data by 90% with only slight performance degradation. To the best of our knowledge, this is the first work aiming to enhance low-resource self-supervised learning in speech processing.
format Preprint
id arxiv_https___arxiv_org_abs_2309_17020
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
Hsu, Po-chun
Elkahky, Ali
Hsu, Wei-Ning
Adi, Yossi
Nguyen, Tu Anh
Copet, Jade
Dupoux, Emmanuel
Lee, Hung-yi
Mohamed, Abdelrahman
Audio and Speech Processing
Sound
Self-supervised learning (SSL) techniques have achieved remarkable results in various speech processing tasks. Nonetheless, a significant challenge remains in reducing the reliance on vast amounts of speech data for pre-training. This paper proposes to address this challenge by leveraging synthetic speech to augment a low-resource pre-training corpus. We construct a high-quality text-to-speech (TTS) system with limited resources using SSL features and generate a large synthetic corpus for pre-training. Experimental results demonstrate that our proposed approach effectively reduces the demand for speech data by 90% with only slight performance degradation. To the best of our knowledge, this is the first work aiming to enhance low-resource self-supervised learning in speech processing.
title Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2309.17020