Deep Learning for Assessment of Oral Reading Fluency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vaidya, Mithilesh, Sahoo, Binaya Kumar, Rao, Preeti
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910467013738496
author Vaidya, Mithilesh
Sahoo, Binaya Kumar
Rao, Preeti
author_facet Vaidya, Mithilesh
Sahoo, Binaya Kumar
Rao, Preeti
contents Reading fluency assessment is a critical component of literacy programmes, serving to guide and monitor early education interventions. Given the resource intensive nature of the exercise when conducted by teachers, the development of automatic tools that can operate on audio recordings of oral reading is attractive as an objective and highly scalable solution. Multiple complex aspects such as accuracy, rate and expressiveness underlie human judgements of reading fluency. In this work, we investigate end-to-end modeling on a training dataset of children's audio recordings of story texts labeled by human experts. The pre-trained wav2vec2.0 model is adopted due its potential to alleviate the challenges from the limited amount of labeled data. We report the performance of a number of system variations on the relevant measures, and also probe the learned embeddings for lexical and acoustic-prosodic features known to be important to the perception of reading fluency.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19426
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Deep Learning for Assessment of Oral Reading Fluency
Vaidya, Mithilesh
Sahoo, Binaya Kumar
Rao, Preeti
Computation and Language
Sound
Audio and Speech Processing
Reading fluency assessment is a critical component of literacy programmes, serving to guide and monitor early education interventions. Given the resource intensive nature of the exercise when conducted by teachers, the development of automatic tools that can operate on audio recordings of oral reading is attractive as an objective and highly scalable solution. Multiple complex aspects such as accuracy, rate and expressiveness underlie human judgements of reading fluency. In this work, we investigate end-to-end modeling on a training dataset of children's audio recordings of story texts labeled by human experts. The pre-trained wav2vec2.0 model is adopted due its potential to alleviate the challenges from the limited amount of labeled data. We report the performance of a number of system variations on the relevant measures, and also probe the learned embeddings for lexical and acoustic-prosodic features known to be important to the perception of reading fluency.
title Deep Learning for Assessment of Oral Reading Fluency
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2405.19426