Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wills, Simone, Bai, Yu, Tejedor-Garcia, Cristian, Cucchiarini, Catia, Strik, Helmer
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914881988460544
author Wills, Simone
Bai, Yu
Tejedor-Garcia, Cristian
Cucchiarini, Catia
Strik, Helmer
author_facet Wills, Simone
Bai, Yu
Tejedor-Garcia, Cristian
Cucchiarini, Catia
Strik, Helmer
contents Voicebots have provided a new avenue for supporting the development of language skills, particularly within the context of second language learning. Voicebots, though, have largely been geared towards native adult speakers. We sought to assess the performance of two state-of-the-art ASR systems, Wav2Vec2.0 and Whisper AI, with a view to developing a voicebot that can support children acquiring a foreign language. We evaluated their performance on read and extemporaneous speech of native and non-native Dutch children. We also investigated the utility of using ASR technology to provide insight into the children's pronunciation and fluency. The results show that recent, pre-trained ASR transformer-based models achieve acceptable performance from which detailed feedback on phoneme pronunciation quality can be extracted, despite the challenging nature of child and non-native speech.
format Preprint
id arxiv_https___arxiv_org_abs_2306_16710
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
Wills, Simone
Bai, Yu
Tejedor-Garcia, Cristian
Cucchiarini, Catia
Strik, Helmer
Computation and Language
Sound
Audio and Speech Processing
Signal Processing
Voicebots have provided a new avenue for supporting the development of language skills, particularly within the context of second language learning. Voicebots, though, have largely been geared towards native adult speakers. We sought to assess the performance of two state-of-the-art ASR systems, Wav2Vec2.0 and Whisper AI, with a view to developing a voicebot that can support children acquiring a foreign language. We evaluated their performance on read and extemporaneous speech of native and non-native Dutch children. We also investigated the utility of using ASR technology to provide insight into the children's pronunciation and fluency. The results show that recent, pre-trained ASR transformer-based models achieve acceptable performance from which detailed feedback on phoneme pronunciation quality can be extracted, despite the challenging nature of child and non-native speech.
title Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
topic Computation and Language
Sound
Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2306.16710