Improving child speech recognition with augmented child-like speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuanyuan, Yue, Zhengjun, Patel, Tanvina, Scharenborg, Odette
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916556654510080
author Zhang, Yuanyuan
Yue, Zhengjun
Patel, Tanvina
Scharenborg, Odette
author_facet Zhang, Yuanyuan
Yue, Zhengjun
Patel, Tanvina
Scharenborg, Odette
contents State-of-the-art ASRs show suboptimal performance for child speech. The scarcity of child speech limits the development of child speech recognition (CSR). Therefore, we studied child-to-child voice conversion (VC) from existing child speakers in the dataset and additional (new) child speakers via monolingual and cross-lingual (Dutch-to-German) VC, respectively. The results showed that cross-lingual child-to-child VC significantly improved child ASR performance. Experiments on the impact of the quantity of child-to-child cross-lingual VC-generated data on fine-tuning (FT) ASR models gave the best results with two-fold augmentation for our FT-Conformer model and FT-Whisper model which reduced WERs with ~3% absolute compared to the baseline, and with six-fold augmentation for the model trained from scratch, which improved by an absolute 3.6% WER. Moreover, using a small amount of "high-quality" VC-generated data achieved similar results to those of our best-FT models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10284
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving child speech recognition with augmented child-like speech
Zhang, Yuanyuan
Yue, Zhengjun
Patel, Tanvina
Scharenborg, Odette
Computation and Language
Sound
Audio and Speech Processing
State-of-the-art ASRs show suboptimal performance for child speech. The scarcity of child speech limits the development of child speech recognition (CSR). Therefore, we studied child-to-child voice conversion (VC) from existing child speakers in the dataset and additional (new) child speakers via monolingual and cross-lingual (Dutch-to-German) VC, respectively. The results showed that cross-lingual child-to-child VC significantly improved child ASR performance. Experiments on the impact of the quantity of child-to-child cross-lingual VC-generated data on fine-tuning (FT) ASR models gave the best results with two-fold augmentation for our FT-Conformer model and FT-Whisper model which reduced WERs with ~3% absolute compared to the baseline, and with six-fold augmentation for the model trained from scratch, which improved by an absolute 3.6% WER. Moreover, using a small amount of "high-quality" VC-generated data achieved similar results to those of our best-FT models.
title Improving child speech recognition with augmented child-like speech
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.10284