Evaluating Logit-Based GOP Scores for Mispronunciation Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Parikh, Aditya Kamlesh, Tejedor-Garcia, Cristian, Cucchiarini, Catia, Strik, Helmer
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914012493512704
author Parikh, Aditya Kamlesh
Tejedor-Garcia, Cristian
Cucchiarini, Catia
Strik, Helmer
author_facet Parikh, Aditya Kamlesh
Tejedor-Garcia, Cristian
Cucchiarini, Catia
Strik, Helmer
contents Pronunciation assessment relies on goodness of pronunciation (GOP) scores, traditionally derived from softmax-based posterior probabilities. However, posterior probabilities may suffer from overconfidence and poor phoneme separation, limiting their effectiveness. This study compares logit-based GOP scores with probability-based GOP scores for mispronunciation detection. We conducted our experiment on two L2 English speech datasets spoken by Dutch and Mandarin speakers, assessing classification performance and correlation with human ratings. Logit-based methods outperform probability-based GOP in classification, but their effectiveness depends on dataset characteristics. The maximum logit GOP shows the strongest alignment with human perception, while a combination of different GOP scores balances probability and logit features. The findings suggest that hybrid GOP methods incorporating uncertainty modeling and phoneme-specific weighting improve pronunciation assessment.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12067
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Logit-Based GOP Scores for Mispronunciation Detection
Parikh, Aditya Kamlesh
Tejedor-Garcia, Cristian
Cucchiarini, Catia
Strik, Helmer
Audio and Speech Processing
Artificial Intelligence
Sound
Pronunciation assessment relies on goodness of pronunciation (GOP) scores, traditionally derived from softmax-based posterior probabilities. However, posterior probabilities may suffer from overconfidence and poor phoneme separation, limiting their effectiveness. This study compares logit-based GOP scores with probability-based GOP scores for mispronunciation detection. We conducted our experiment on two L2 English speech datasets spoken by Dutch and Mandarin speakers, assessing classification performance and correlation with human ratings. Logit-based methods outperform probability-based GOP in classification, but their effectiveness depends on dataset characteristics. The maximum logit GOP shows the strongest alignment with human perception, while a combination of different GOP scores balances probability and logit features. The findings suggest that hybrid GOP methods incorporating uncertainty modeling and phoneme-specific weighting improve pronunciation assessment.
title Evaluating Logit-Based GOP Scores for Mispronunciation Detection
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2506.12067