MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Yu-Wen, Yu, Zhou, Hirschberg, Julia
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909216833273856
author Chen, Yu-Wen
Yu, Zhou
Hirschberg, Julia
author_facet Chen, Yu-Wen
Yu, Zhou
Hirschberg, Julia
contents Pronunciation assessment models designed for open response scenarios enable users to practice language skills in a manner similar to real-life communication. However, previous open-response pronunciation assessment models have predominantly focused on a single pronunciation task, such as sentence-level accuracy, rather than offering a comprehensive assessment in various aspects. We propose MultiPA, a Multitask Pronunciation Assessment model that provides sentence-level accuracy, fluency, prosody, and word-level accuracy assessment for open responses. We examined the correlation between different pronunciation tasks and showed the benefits of multi-task learning. Our model reached the state-of-the-art performance on existing in-domain data sets and effectively generalized to an out-of-domain dataset that we newly collected. The experimental results demonstrate the practical utility of our model in real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2308_12490
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
Chen, Yu-Wen
Yu, Zhou
Hirschberg, Julia
Computation and Language
Sound
Audio and Speech Processing
Pronunciation assessment models designed for open response scenarios enable users to practice language skills in a manner similar to real-life communication. However, previous open-response pronunciation assessment models have predominantly focused on a single pronunciation task, such as sentence-level accuracy, rather than offering a comprehensive assessment in various aspects. We propose MultiPA, a Multitask Pronunciation Assessment model that provides sentence-level accuracy, fluency, prosody, and word-level accuracy assessment for open responses. We examined the correlation between different pronunciation tasks and showed the benefits of multi-task learning. Our model reached the state-of-the-art performance on existing in-domain data sets and effectively generalized to an out-of-domain dataset that we newly collected. The experimental results demonstrate the practical utility of our model in real-world applications.
title MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2308.12490