Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908279299375104 |
|---|---|
| author | Choi, Kwanghee Yeo, Eunjung Chang, Kalvin Watanabe, Shinji Mortensen, David |
| author_facet | Choi, Kwanghee Yeo, Eunjung Chang, Kalvin Watanabe, Shinji Mortensen, David |
| contents | Allophony refers to the variation in the phonetic realization of a phoneme based on its phonetic environment. Modeling allophones is crucial for atypical pronunciation assessment, which involves distinguishing atypical from typical pronunciations. However, recent phoneme classifier-based approaches often simplify this by treating various realizations as a single phoneme, bypassing the complexity of modeling allophonic variation. Motivated by the acoustic modeling capabilities of frozen self-supervised speech model (S3M) features, we propose MixGoP, a novel approach that leverages Gaussian mixture models to model phoneme distributions with multiple subclusters. Our experiments show that MixGoP achieves state-of-the-art performance across four out of five datasets, including dysarthric and non-native speech. Our analysis further suggests that S3M features capture allophonic variation more effectively than MFCCs and Mel spectrograms, highlighting the benefits of integrating MixGoP with S3M features. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_07029 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment Choi, Kwanghee Yeo, Eunjung Chang, Kalvin Watanabe, Shinji Mortensen, David Computation and Language Artificial Intelligence Machine Learning Audio and Speech Processing Allophony refers to the variation in the phonetic realization of a phoneme based on its phonetic environment. Modeling allophones is crucial for atypical pronunciation assessment, which involves distinguishing atypical from typical pronunciations. However, recent phoneme classifier-based approaches often simplify this by treating various realizations as a single phoneme, bypassing the complexity of modeling allophonic variation. Motivated by the acoustic modeling capabilities of frozen self-supervised speech model (S3M) features, we propose MixGoP, a novel approach that leverages Gaussian mixture models to model phoneme distributions with multiple subclusters. Our experiments show that MixGoP achieves state-of-the-art performance across four out of five datasets, including dysarthric and non-native speech. Our analysis further suggests that S3M features capture allophonic variation more effectively than MFCCs and Mel spectrograms, highlighting the benefits of integrating MixGoP with S3M features. |
| title | Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment |
| topic | Computation and Language Artificial Intelligence Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2502.07029 |