Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917683991150592 |
|---|---|
| author | Hsu, Po-chun Elkahky, Ali Hsu, Wei-Ning Adi, Yossi Nguyen, Tu Anh Copet, Jade Dupoux, Emmanuel Lee, Hung-yi Mohamed, Abdelrahman |
| author_facet | Hsu, Po-chun Elkahky, Ali Hsu, Wei-Ning Adi, Yossi Nguyen, Tu Anh Copet, Jade Dupoux, Emmanuel Lee, Hung-yi Mohamed, Abdelrahman |
| contents | Self-supervised learning (SSL) techniques have achieved remarkable results in various speech processing tasks. Nonetheless, a significant challenge remains in reducing the reliance on vast amounts of speech data for pre-training. This paper proposes to address this challenge by leveraging synthetic speech to augment a low-resource pre-training corpus. We construct a high-quality text-to-speech (TTS) system with limited resources using SSL features and generate a large synthetic corpus for pre-training. Experimental results demonstrate that our proposed approach effectively reduces the demand for speech data by 90% with only slight performance degradation. To the best of our knowledge, this is the first work aiming to enhance low-resource self-supervised learning in speech processing. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2309_17020 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Low-Resource Self-Supervised Learning with SSL-Enhanced TTS Hsu, Po-chun Elkahky, Ali Hsu, Wei-Ning Adi, Yossi Nguyen, Tu Anh Copet, Jade Dupoux, Emmanuel Lee, Hung-yi Mohamed, Abdelrahman Audio and Speech Processing Sound Self-supervised learning (SSL) techniques have achieved remarkable results in various speech processing tasks. Nonetheless, a significant challenge remains in reducing the reliance on vast amounts of speech data for pre-training. This paper proposes to address this challenge by leveraging synthetic speech to augment a low-resource pre-training corpus. We construct a high-quality text-to-speech (TTS) system with limited resources using SSL features and generate a large synthetic corpus for pre-training. Experimental results demonstrate that our proposed approach effectively reduces the demand for speech data by 90% with only slight performance degradation. To the best of our knowledge, this is the first work aiming to enhance low-resource self-supervised learning in speech processing. |
| title | Low-Resource Self-Supervised Learning with SSL-Enhanced TTS |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2309.17020 |