Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge
Fuente:
arXiv
Salvato in:
| Autori principali: | , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916930988802048 |
|---|---|
| author | Žavoronkov, Aleksei Alumäe, Tanel |
| author_facet | Žavoronkov, Aleksei Alumäe, Tanel |
| contents | This paper presents an analysis of three end-to-end models developed for the NOCASA 2025 Challenge, aimed at automatic word-level pronunciation assessment for children learning Norwegian as a second language. Our models include an encoder-decoder Siamese architecture (E2E-R), a prefix-tuned direct classification model leveraging pretrained wav2vec2.0 representations, and a novel model integrating alignment-free goodness-of-pronunciation (GOP) features computed via CTC. We introduce a weighted ordinal cross-entropy loss tailored for optimizing metrics such as unweighted average recall and mean absolute error. Among the explored methods, our GOP-CTC-based model achieved the highest performance, substantially surpassing challenge baselines and attaining top leaderboard scores. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_03256 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge Žavoronkov, Aleksei Alumäe, Tanel Computation and Language Sound Audio and Speech Processing This paper presents an analysis of three end-to-end models developed for the NOCASA 2025 Challenge, aimed at automatic word-level pronunciation assessment for children learning Norwegian as a second language. Our models include an encoder-decoder Siamese architecture (E2E-R), a prefix-tuned direct classification model leveraging pretrained wav2vec2.0 representations, and a novel model integrating alignment-free goodness-of-pronunciation (GOP) features computed via CTC. We introduce a weighted ordinal cross-entropy loss tailored for optimizing metrics such as unweighted average recall and mean absolute error. Among the explored methods, our GOP-CTC-based model achieved the highest performance, substantially surpassing challenge baselines and attaining top leaderboard scores. |
| title | Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2509.03256 |