An End-to-End Approach for Child Reading Assessment in the Xhosa Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chevtchenko, Sergio, Navas, Nikhil, Vale, Rafaella, Ubaudi, Franco, Lucwaba, Sipumelele, Ardington, Cally, Afshar, Soheil, Antoniou, Mark, Afshar, Saeed
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918042628259840
author Chevtchenko, Sergio
Navas, Nikhil
Vale, Rafaella
Ubaudi, Franco
Lucwaba, Sipumelele
Ardington, Cally
Afshar, Soheil
Antoniou, Mark
Afshar, Saeed
author_facet Chevtchenko, Sergio
Navas, Nikhil
Vale, Rafaella
Ubaudi, Franco
Lucwaba, Sipumelele
Ardington, Cally
Afshar, Soheil
Antoniou, Mark
Afshar, Saeed
contents Child literacy is a strong predictor of life outcomes at the subsequent stages of an individual's life. This points to a need for targeted interventions in vulnerable low and middle income populations to help bridge the gap between literacy levels in these regions and high income ones. In this effort, reading assessments provide an important tool to measure the effectiveness of these programs and AI can be a reliable and economical tool to support educators with this task. Developing accurate automatic reading assessment systems for child speech in low-resource languages poses significant challenges due to limited data and the unique acoustic properties of children's voices. This study focuses on Xhosa, a language spoken in South Africa, to advance child speech recognition capabilities. We present a novel dataset composed of child speech samples in Xhosa. The dataset is available upon request and contains ten words and letters, which are part of the Early Grade Reading Assessment (EGRA) system. Each recording is labeled with an online and cost-effective approach by multiple markers and a subsample is validated by an independent EGRA reviewer. This dataset is evaluated with three fine-tuned state-of-the-art end-to-end models: wav2vec 2.0, HuBERT, and Whisper. The results indicate that the performance of these models can be significantly influenced by the amount and balancing of the available training data, which is fundamental for cost-effective large dataset collection. Furthermore, our experiments indicate that the wav2vec 2.0 performance is improved by training on multiple classes at a time, even when the number of available samples is constrained.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17371
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An End-to-End Approach for Child Reading Assessment in the Xhosa Language
Chevtchenko, Sergio
Navas, Nikhil
Vale, Rafaella
Ubaudi, Franco
Lucwaba, Sipumelele
Ardington, Cally
Afshar, Soheil
Antoniou, Mark
Afshar, Saeed
Machine Learning
Computation and Language
Child literacy is a strong predictor of life outcomes at the subsequent stages of an individual's life. This points to a need for targeted interventions in vulnerable low and middle income populations to help bridge the gap between literacy levels in these regions and high income ones. In this effort, reading assessments provide an important tool to measure the effectiveness of these programs and AI can be a reliable and economical tool to support educators with this task. Developing accurate automatic reading assessment systems for child speech in low-resource languages poses significant challenges due to limited data and the unique acoustic properties of children's voices. This study focuses on Xhosa, a language spoken in South Africa, to advance child speech recognition capabilities. We present a novel dataset composed of child speech samples in Xhosa. The dataset is available upon request and contains ten words and letters, which are part of the Early Grade Reading Assessment (EGRA) system. Each recording is labeled with an online and cost-effective approach by multiple markers and a subsample is validated by an independent EGRA reviewer. This dataset is evaluated with three fine-tuned state-of-the-art end-to-end models: wav2vec 2.0, HuBERT, and Whisper. The results indicate that the performance of these models can be significantly influenced by the amount and balancing of the available training data, which is fundamental for cost-effective large dataset collection. Furthermore, our experiments indicate that the wav2vec 2.0 performance is improved by training on multiple classes at a time, even when the number of available samples is constrained.
title An End-to-End Approach for Child Reading Assessment in the Xhosa Language
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2505.17371