Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909696232783872 |
|---|---|
| author | Zhou, Xuanru Lian, Jiachen Cho, Cheol Jun Prabhune, Tejas Li, Shuhe Li, William Ortiz, Rodrigo Ezzes, Zoe Vonk, Jet Morin, Brittany Bogley, Rian Wauters, Lisa Miller, Zachary Gorno-Tempini, Maria Anumanchipalli, Gopala |
| author_facet | Zhou, Xuanru Lian, Jiachen Cho, Cheol Jun Prabhune, Tejas Li, Shuhe Li, William Ortiz, Rodrigo Ezzes, Zoe Vonk, Jet Morin, Brittany Bogley, Rian Wauters, Lisa Miller, Zachary Gorno-Tempini, Maria Anumanchipalli, Gopala |
| contents | Phonetic error detection, a core subtask of automatic pronunciation assessment, identifies pronunciation deviations at the phoneme level. Speech variability from accents and dysfluencies challenges accurate phoneme recognition, with current models failing to capture these discrepancies effectively. We propose a verbatim phoneme recognition framework using multi-task training with novel phoneme similarity modeling that transcribes what speakers actually say rather than what they're supposed to say. We develop and open-source \textit{VCTK-accent}, a simulated dataset containing phonetic errors, and propose two novel metrics for assessing pronunciation differences. Our work establishes a new benchmark for phonetic error detection. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_14346 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling Zhou, Xuanru Lian, Jiachen Cho, Cheol Jun Prabhune, Tejas Li, Shuhe Li, William Ortiz, Rodrigo Ezzes, Zoe Vonk, Jet Morin, Brittany Bogley, Rian Wauters, Lisa Miller, Zachary Gorno-Tempini, Maria Anumanchipalli, Gopala Audio and Speech Processing Sound Phonetic error detection, a core subtask of automatic pronunciation assessment, identifies pronunciation deviations at the phoneme level. Speech variability from accents and dysfluencies challenges accurate phoneme recognition, with current models failing to capture these discrepancies effectively. We propose a verbatim phoneme recognition framework using multi-task training with novel phoneme similarity modeling that transcribes what speakers actually say rather than what they're supposed to say. We develop and open-source \textit{VCTK-accent}, a simulated dataset containing phonetic errors, and propose two novel metrics for assessing pronunciation differences. Our work establishes a new benchmark for phonetic error detection. |
| title | Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2507.14346 |