Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Žavoronkov, Aleksei, Alumäe, Tanel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916930988802048
author Žavoronkov, Aleksei
Alumäe, Tanel
author_facet Žavoronkov, Aleksei
Alumäe, Tanel
contents This paper presents an analysis of three end-to-end models developed for the NOCASA 2025 Challenge, aimed at automatic word-level pronunciation assessment for children learning Norwegian as a second language. Our models include an encoder-decoder Siamese architecture (E2E-R), a prefix-tuned direct classification model leveraging pretrained wav2vec2.0 representations, and a novel model integrating alignment-free goodness-of-pronunciation (GOP) features computed via CTC. We introduce a weighted ordinal cross-entropy loss tailored for optimizing metrics such as unweighted average recall and mean absolute error. Among the explored methods, our GOP-CTC-based model achieved the highest performance, substantially surpassing challenge baselines and attaining top leaderboard scores.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03256
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge
Žavoronkov, Aleksei
Alumäe, Tanel
Computation and Language
Sound
Audio and Speech Processing
This paper presents an analysis of three end-to-end models developed for the NOCASA 2025 Challenge, aimed at automatic word-level pronunciation assessment for children learning Norwegian as a second language. Our models include an encoder-decoder Siamese architecture (E2E-R), a prefix-tuned direct classification model leveraging pretrained wav2vec2.0 representations, and a novel model integrating alignment-free goodness-of-pronunciation (GOP) features computed via CTC. We introduce a weighted ordinal cross-entropy loss tailored for optimizing metrics such as unweighted average recall and mean absolute error. Among the explored methods, our GOP-CTC-based model achieved the highest performance, substantially surpassing challenge baselines and attaining top leaderboard scores.
title Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.03256