MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909216833273856 |
|---|---|
| author | Chen, Yu-Wen Yu, Zhou Hirschberg, Julia |
| author_facet | Chen, Yu-Wen Yu, Zhou Hirschberg, Julia |
| contents | Pronunciation assessment models designed for open response scenarios enable users to practice language skills in a manner similar to real-life communication. However, previous open-response pronunciation assessment models have predominantly focused on a single pronunciation task, such as sentence-level accuracy, rather than offering a comprehensive assessment in various aspects. We propose MultiPA, a Multitask Pronunciation Assessment model that provides sentence-level accuracy, fluency, prosody, and word-level accuracy assessment for open responses. We examined the correlation between different pronunciation tasks and showed the benefits of multi-task learning. Our model reached the state-of-the-art performance on existing in-domain data sets and effectively generalized to an out-of-domain dataset that we newly collected. The experimental results demonstrate the practical utility of our model in real-world applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2308_12490 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios Chen, Yu-Wen Yu, Zhou Hirschberg, Julia Computation and Language Sound Audio and Speech Processing Pronunciation assessment models designed for open response scenarios enable users to practice language skills in a manner similar to real-life communication. However, previous open-response pronunciation assessment models have predominantly focused on a single pronunciation task, such as sentence-level accuracy, rather than offering a comprehensive assessment in various aspects. We propose MultiPA, a Multitask Pronunciation Assessment model that provides sentence-level accuracy, fluency, prosody, and word-level accuracy assessment for open responses. We examined the correlation between different pronunciation tasks and showed the benefits of multi-task learning. Our model reached the state-of-the-art performance on existing in-domain data sets and effectively generalized to an out-of-domain dataset that we newly collected. The experimental results demonstrate the practical utility of our model in real-world applications. |
| title | MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2308.12490 |