HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Bi-Cheng, Wang, Hsin-Wei, Chao, Fu-An, Lo, Tien-Hong, Hsu, Yung-Chang, Chen, Berlin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918231875256320
author Yan, Bi-Cheng
Wang, Hsin-Wei
Chao, Fu-An
Lo, Tien-Hong
Hsu, Yung-Chang
Chen, Berlin
author_facet Yan, Bi-Cheng
Wang, Hsin-Wei
Chao, Fu-An
Lo, Tien-Hong
Hsu, Yung-Chang
Chen, Berlin
contents Automatic pronunciation assessment (APA) seeks to quantify a second language (L2) learner's pronunciation proficiency in a target language by offering timely and fine-grained diagnostic feedback. Most existing efforts on APA have predominantly concentrated on highly constrained reading-aloud tasks (where learners are prompted to read a reference text aloud); however, assessing pronunciation quality in unscripted speech (or free-speaking scenarios) remains relatively underexplored. In light of this, we first propose HiPPO, a hierarchical pronunciation assessment model tailored for spoken languages, which evaluates an L2 learner's oral proficiency at multiple linguistic levels based solely on the speech uttered by the learner. To improve the overall accuracy of assessment, a contrastive ordinal regularizer and a curriculum learning strategy are introduced for model training. The former aims to generate score-discriminative features by exploiting the ordinal nature of regression targets, while the latter gradually ramps up the training complexity to facilitate the assessment task that takes unscripted speech as input. Experiments conducted on the Speechocean762 benchmark dataset validates the feasibility and superiority of our method in relation to several cutting-edge baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04964
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
Yan, Bi-Cheng
Wang, Hsin-Wei
Chao, Fu-An
Lo, Tien-Hong
Hsu, Yung-Chang
Chen, Berlin
Audio and Speech Processing
Automatic pronunciation assessment (APA) seeks to quantify a second language (L2) learner's pronunciation proficiency in a target language by offering timely and fine-grained diagnostic feedback. Most existing efforts on APA have predominantly concentrated on highly constrained reading-aloud tasks (where learners are prompted to read a reference text aloud); however, assessing pronunciation quality in unscripted speech (or free-speaking scenarios) remains relatively underexplored. In light of this, we first propose HiPPO, a hierarchical pronunciation assessment model tailored for spoken languages, which evaluates an L2 learner's oral proficiency at multiple linguistic levels based solely on the speech uttered by the learner. To improve the overall accuracy of assessment, a contrastive ordinal regularizer and a curriculum learning strategy are introduced for model training. The former aims to generate score-discriminative features by exploiting the ordinal nature of regression targets, while the latter gradually ramps up the training complexity to facilitate the assessment task that takes unscripted speech as input. Experiments conducted on the Speechocean762 benchmark dataset validates the feasibility and superiority of our method in relation to several cutting-edge baselines.
title HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
topic Audio and Speech Processing
url https://arxiv.org/abs/2512.04964