Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Choi, Kwanghee, Yeo, Eunjung, Chang, Kalvin, Watanabe, Shinji, Mortensen, David
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908279299375104
author Choi, Kwanghee
Yeo, Eunjung
Chang, Kalvin
Watanabe, Shinji
Mortensen, David
author_facet Choi, Kwanghee
Yeo, Eunjung
Chang, Kalvin
Watanabe, Shinji
Mortensen, David
contents Allophony refers to the variation in the phonetic realization of a phoneme based on its phonetic environment. Modeling allophones is crucial for atypical pronunciation assessment, which involves distinguishing atypical from typical pronunciations. However, recent phoneme classifier-based approaches often simplify this by treating various realizations as a single phoneme, bypassing the complexity of modeling allophonic variation. Motivated by the acoustic modeling capabilities of frozen self-supervised speech model (S3M) features, we propose MixGoP, a novel approach that leverages Gaussian mixture models to model phoneme distributions with multiple subclusters. Our experiments show that MixGoP achieves state-of-the-art performance across four out of five datasets, including dysarthric and non-native speech. Our analysis further suggests that S3M features capture allophonic variation more effectively than MFCCs and Mel spectrograms, highlighting the benefits of integrating MixGoP with S3M features.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07029
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
Choi, Kwanghee
Yeo, Eunjung
Chang, Kalvin
Watanabe, Shinji
Mortensen, David
Computation and Language
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Allophony refers to the variation in the phonetic realization of a phoneme based on its phonetic environment. Modeling allophones is crucial for atypical pronunciation assessment, which involves distinguishing atypical from typical pronunciations. However, recent phoneme classifier-based approaches often simplify this by treating various realizations as a single phoneme, bypassing the complexity of modeling allophonic variation. Motivated by the acoustic modeling capabilities of frozen self-supervised speech model (S3M) features, we propose MixGoP, a novel approach that leverages Gaussian mixture models to model phoneme distributions with multiple subclusters. Our experiments show that MixGoP achieves state-of-the-art performance across four out of five datasets, including dysarthric and non-native speech. Our analysis further suggests that S3M features capture allophonic variation more effectively than MFCCs and Mel spectrograms, highlighting the benefits of integrating MixGoP with S3M features.
title Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
topic Computation and Language
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2502.07029