AMPS: ASR with Multimodal Paraphrase Supervision
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908322989342720 |
|---|---|
| author | Gupta, Abhishek Parulekar, Amruta Chattopadhyay, Sameep Jyothi, Preethi |
| author_facet | Gupta, Abhishek Parulekar, Amruta Chattopadhyay, Sameep Jyothi, Preethi |
| contents | Spontaneous or conversational multilingual speech presents many challenges for state-of-the-art automatic speech recognition (ASR) systems. In this work, we present a new technique AMPS that augments a multilingual multimodal ASR system with paraphrase-based supervision for improved conversational ASR in multiple languages, including Hindi, Marathi, Malayalam, Kannada, and Nyanja. We use paraphrases of the reference transcriptions as additional supervision while training the multimodal ASR model and selectively invoke this paraphrase objective for utterances with poor ASR performance. Using AMPS with a state-of-the-art multimodal model SeamlessM4T, we obtain significant relative reductions in word error rates (WERs) of up to 5%. We present detailed analyses of our system using both objective and human evaluation metrics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_18368 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | AMPS: ASR with Multimodal Paraphrase Supervision Gupta, Abhishek Parulekar, Amruta Chattopadhyay, Sameep Jyothi, Preethi Computation and Language Artificial Intelligence Machine Learning Audio and Speech Processing Spontaneous or conversational multilingual speech presents many challenges for state-of-the-art automatic speech recognition (ASR) systems. In this work, we present a new technique AMPS that augments a multilingual multimodal ASR system with paraphrase-based supervision for improved conversational ASR in multiple languages, including Hindi, Marathi, Malayalam, Kannada, and Nyanja. We use paraphrases of the reference transcriptions as additional supervision while training the multimodal ASR model and selectively invoke this paraphrase objective for utterances with poor ASR performance. Using AMPS with a state-of-the-art multimodal model SeamlessM4T, we obtain significant relative reductions in word error rates (WERs) of up to 5%. We present detailed analyses of our system using both objective and human evaluation metrics. |
| title | AMPS: ASR with Multimodal Paraphrase Supervision |
| topic | Computation and Language Artificial Intelligence Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2411.18368 |