Automated evaluation of children's speech fluency for low-resource languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Bowen, Latiff, Nur Afiqah Abdul, Kan, Justin, Tong, Rong, Soh, Donny, Miao, Xiaoxiao, McLoughlin, Ian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912665496977408
author Zhang, Bowen
Latiff, Nur Afiqah Abdul
Kan, Justin
Tong, Rong
Soh, Donny
Miao, Xiaoxiao
McLoughlin, Ian
author_facet Zhang, Bowen
Latiff, Nur Afiqah Abdul
Kan, Justin
Tong, Rong
Soh, Donny
Miao, Xiaoxiao
McLoughlin, Ian
contents Assessment of children's speaking fluency in education is well researched for majority languages, but remains highly challenging for low resource languages. This paper proposes a system to automatically assess fluency by combining a fine-tuned multilingual ASR model, an objective metrics extraction stage, and a generative pre-trained transformer (GPT) network. The objective metrics include phonetic and word error rates, speech rate, and speech-pause duration ratio. These are interpreted by a GPT-based classifier guided by a small set of human-evaluated ground truth examples, to score fluency. We evaluate the proposed system on a dataset of children's speech in two low-resource languages, Tamil and Malay and compare the classification performance against Random Forest and XGBoost, as well as using ChatGPT-4o to predict fluency directly from speech input. Results demonstrate that the proposed approach achieves significantly higher accuracy than multimodal GPT or other methods.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19671
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automated evaluation of children's speech fluency for low-resource languages
Zhang, Bowen
Latiff, Nur Afiqah Abdul
Kan, Justin
Tong, Rong
Soh, Donny
Miao, Xiaoxiao
McLoughlin, Ian
Sound
Artificial Intelligence
Audio and Speech Processing
Assessment of children's speaking fluency in education is well researched for majority languages, but remains highly challenging for low resource languages. This paper proposes a system to automatically assess fluency by combining a fine-tuned multilingual ASR model, an objective metrics extraction stage, and a generative pre-trained transformer (GPT) network. The objective metrics include phonetic and word error rates, speech rate, and speech-pause duration ratio. These are interpreted by a GPT-based classifier guided by a small set of human-evaluated ground truth examples, to score fluency. We evaluate the proposed system on a dataset of children's speech in two low-resource languages, Tamil and Malay and compare the classification performance against Random Forest and XGBoost, as well as using ChatGPT-4o to predict fluency directly from speech input. Results demonstrate that the proposed approach achieves significantly higher accuracy than multimodal GPT or other methods.
title Automated evaluation of children's speech fluency for low-resource languages
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.19671