Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Phan, Nhan, Kuronen, Mikko, Kautonen, Maria, Ullakonoja, Riikka, von Zansen, Anna, Getman, Yaroslav, Voskoboinik, Ekaterina, Grósz, Tamás, Kurimo, Mikko
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913998006386688
author Phan, Nhan
Kuronen, Mikko
Kautonen, Maria
Ullakonoja, Riikka
von Zansen, Anna
Getman, Yaroslav
Voskoboinik, Ekaterina
Grósz, Tamás
Kurimo, Mikko
author_facet Phan, Nhan
Kuronen, Mikko
Kautonen, Maria
Ullakonoja, Riikka
von Zansen, Anna
Getman, Yaroslav
Voskoboinik, Ekaterina
Grósz, Tamás
Kurimo, Mikko
contents Mispronunciation detection (MD) models are the cornerstones of many language learning applications. Unfortunately, most systems are built for English and other major languages, while low-resourced language varieties, such as Finland Swedish (FS), lack such tools. In this paper, we introduce our MD model for FS, trained on 89 hours of first language (L1) speakers' spontaneous speech and tested on 33 minutes of L2 transcribed read-aloud speech. We trained a multilingual wav2vec 2.0 model with entropy regularization, followed by temperature scaling and top-k normalization after the inference to better adapt it for MD. The main novelty of our method lies in its simplicity, requiring minimal L2 data. The process is also language-independent, making it suitable for other low-resource languages. Our proposed algorithm allows us to balance Recall (43.2%) and Precision (29.8%), compared with the baseline model's Recall (77.5%) and Precision (17.6%).
format Preprint
id arxiv_https___arxiv_org_abs_2506_01156
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish
Phan, Nhan
Kuronen, Mikko
Kautonen, Maria
Ullakonoja, Riikka
von Zansen, Anna
Getman, Yaroslav
Voskoboinik, Ekaterina
Grósz, Tamás
Kurimo, Mikko
Computation and Language
Sound
Audio and Speech Processing
Mispronunciation detection (MD) models are the cornerstones of many language learning applications. Unfortunately, most systems are built for English and other major languages, while low-resourced language varieties, such as Finland Swedish (FS), lack such tools. In this paper, we introduce our MD model for FS, trained on 89 hours of first language (L1) speakers' spontaneous speech and tested on 33 minutes of L2 transcribed read-aloud speech. We trained a multilingual wav2vec 2.0 model with entropy regularization, followed by temperature scaling and top-k normalization after the inference to better adapt it for MD. The main novelty of our method lies in its simplicity, requiring minimal L2 data. The process is also language-independent, making it suitable for other low-resource languages. Our proposed algorithm allows us to balance Recall (43.2%) and Precision (29.8%), compared with the baseline model's Recall (77.5%) and Precision (17.6%).
title Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.01156