Causal Structure Discovery for Error Diagnostics of Children's ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Vishwanath Pratap, Sahidullah, Md., Kinnunen, Tomi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915315817906176
author Singh, Vishwanath Pratap
Sahidullah, Md.
Kinnunen, Tomi
author_facet Singh, Vishwanath Pratap
Sahidullah, Md.
Kinnunen, Tomi
contents Children's automatic speech recognition (ASR) often underperforms compared to that of adults due to a confluence of interdependent factors: physiological (e.g., smaller vocal tracts), cognitive (e.g., underdeveloped pronunciation), and extrinsic (e.g., vocabulary limitations, background noise). Existing analysis methods examine the impact of these factors in isolation, neglecting interdependencies-such as age affecting ASR accuracy both directly and indirectly via pronunciation skills. In this paper, we introduce a causal structure discovery to unravel these interdependent relationships among physiology, cognition, extrinsic factors, and ASR errors. Then, we employ causal quantification to measure each factor's impact on children's ASR. We extend the analysis to fine-tuned models to identify which factors are mitigated by fine-tuning and which remain largely unaffected. Experiments on Whisper and Wav2Vec2.0 demonstrate the generalizability of our findings across different ASR systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00402
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Causal Structure Discovery for Error Diagnostics of Children's ASR
Singh, Vishwanath Pratap
Sahidullah, Md.
Kinnunen, Tomi
Computation and Language
Sound
Audio and Speech Processing
Children's automatic speech recognition (ASR) often underperforms compared to that of adults due to a confluence of interdependent factors: physiological (e.g., smaller vocal tracts), cognitive (e.g., underdeveloped pronunciation), and extrinsic (e.g., vocabulary limitations, background noise). Existing analysis methods examine the impact of these factors in isolation, neglecting interdependencies-such as age affecting ASR accuracy both directly and indirectly via pronunciation skills. In this paper, we introduce a causal structure discovery to unravel these interdependent relationships among physiology, cognition, extrinsic factors, and ASR errors. Then, we employ causal quantification to measure each factor's impact on children's ASR. We extend the analysis to fine-tuned models to identify which factors are mitigated by fine-tuning and which remain largely unaffected. Experiments on Whisper and Wav2Vec2.0 demonstrate the generalizability of our findings across different ASR systems.
title Causal Structure Discovery for Error Diagnostics of Children's ASR
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.00402