AMPS: ASR with Multimodal Paraphrase Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Abhishek, Parulekar, Amruta, Chattopadhyay, Sameep, Jyothi, Preethi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908322989342720
author Gupta, Abhishek
Parulekar, Amruta
Chattopadhyay, Sameep
Jyothi, Preethi
author_facet Gupta, Abhishek
Parulekar, Amruta
Chattopadhyay, Sameep
Jyothi, Preethi
contents Spontaneous or conversational multilingual speech presents many challenges for state-of-the-art automatic speech recognition (ASR) systems. In this work, we present a new technique AMPS that augments a multilingual multimodal ASR system with paraphrase-based supervision for improved conversational ASR in multiple languages, including Hindi, Marathi, Malayalam, Kannada, and Nyanja. We use paraphrases of the reference transcriptions as additional supervision while training the multimodal ASR model and selectively invoke this paraphrase objective for utterances with poor ASR performance. Using AMPS with a state-of-the-art multimodal model SeamlessM4T, we obtain significant relative reductions in word error rates (WERs) of up to 5%. We present detailed analyses of our system using both objective and human evaluation metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18368
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AMPS: ASR with Multimodal Paraphrase Supervision
Gupta, Abhishek
Parulekar, Amruta
Chattopadhyay, Sameep
Jyothi, Preethi
Computation and Language
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Spontaneous or conversational multilingual speech presents many challenges for state-of-the-art automatic speech recognition (ASR) systems. In this work, we present a new technique AMPS that augments a multilingual multimodal ASR system with paraphrase-based supervision for improved conversational ASR in multiple languages, including Hindi, Marathi, Malayalam, Kannada, and Nyanja. We use paraphrases of the reference transcriptions as additional supervision while training the multimodal ASR model and selectively invoke this paraphrase objective for utterances with poor ASR performance. Using AMPS with a state-of-the-art multimodal model SeamlessM4T, we obtain significant relative reductions in word error rates (WERs) of up to 5%. We present detailed analyses of our system using both objective and human evaluation metrics.
title AMPS: ASR with Multimodal Paraphrase Supervision
topic Computation and Language
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2411.18368