Utterance-Level Methods for Identifying Reliable ASR-Output for Child Speech

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lathouwers, Gus, Gao, Lingyun, Cucchiarini, Catia, Strik, Helmer
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908985747046400
author Lathouwers, Gus
Gao, Lingyun
Cucchiarini, Catia
Strik, Helmer
author_facet Lathouwers, Gus
Gao, Lingyun
Cucchiarini, Catia
Strik, Helmer
contents Automatic Speech Recognition (ASR) is increasingly used in applications involving child speech, such as language learning and literacy acquisition. However, the effectiveness of such applications is limited by high ASR error rates. The negative effects can be mitigated by identifying in advance which ASR-outputs are reliable. This work aims to develop two novel approaches for selecting reliable ASR-output at the utterance level, one for selecting reliable read speech and one for dialogue speech material. Evaluations were done on an English and a Dutch dataset, each with a baseline and finetuned model. The results show that utterance-level selection methods for identifying reliably transcribed speech recordings have high precision for the best strategy (P > 97.4) for both read speech and dialogue material, for both languages. Using the current optimal strategy allows 21.0% to 55.9% of dialogue/read speech datasets to be automatically selected with low (UER of < 2.6) error rates.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19801
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Utterance-Level Methods for Identifying Reliable ASR-Output for Child Speech
Lathouwers, Gus
Gao, Lingyun
Cucchiarini, Catia
Strik, Helmer
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Automatic Speech Recognition (ASR) is increasingly used in applications involving child speech, such as language learning and literacy acquisition. However, the effectiveness of such applications is limited by high ASR error rates. The negative effects can be mitigated by identifying in advance which ASR-outputs are reliable. This work aims to develop two novel approaches for selecting reliable ASR-output at the utterance level, one for selecting reliable read speech and one for dialogue speech material. Evaluations were done on an English and a Dutch dataset, each with a baseline and finetuned model. The results show that utterance-level selection methods for identifying reliably transcribed speech recordings have high precision for the best strategy (P > 97.4) for both read speech and dialogue material, for both languages. Using the current optimal strategy allows 21.0% to 55.9% of dialogue/read speech datasets to be automatically selected with low (UER of < 2.6) error rates.
title Utterance-Level Methods for Identifying Reliable ASR-Output for Child Speech
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.19801