Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Kun, Li, Bei, Lv, Hang, Lu, Quan, Jiang, Ning, Xie, Lei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914772732084224
author Wei, Kun
Li, Bei
Lv, Hang
Lu, Quan
Jiang, Ning
Xie, Lei
author_facet Wei, Kun
Li, Bei
Lv, Hang
Lu, Quan
Jiang, Ning
Xie, Lei
contents Automatic Speech Recognition (ASR) in conversational settings presents unique challenges, including extracting relevant contextual information from previous conversational turns. Due to irrelevant content, error propagation, and redundancy, existing methods struggle to extract longer and more effective contexts. To address this issue, we introduce a novel conversational ASR system, extending the Conformer encoder-decoder model with cross-modal conversational representation. Our approach leverages a cross-modal extractor that combines pre-trained speech and text models through a specialized encoder and a modal-level mask input. This enables the extraction of richer historical speech context without explicit error propagation. We also incorporate conditional latent variational modules to learn conversational level attributes such as role preference and topic coherence. By introducing both cross-modal and conversational representations into the decoder, our model retains context over longer sentences without information loss, achieving relative accuracy improvements of 8.8% and 23% on Mandarin conversation datasets HKUST and MagicData-RAMC, respectively, compared to the standard Conformer model.
format Preprint
id arxiv_https___arxiv_org_abs_2310_14278
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
Wei, Kun
Li, Bei
Lv, Hang
Lu, Quan
Jiang, Ning
Xie, Lei
Sound
Computation and Language
Audio and Speech Processing
Automatic Speech Recognition (ASR) in conversational settings presents unique challenges, including extracting relevant contextual information from previous conversational turns. Due to irrelevant content, error propagation, and redundancy, existing methods struggle to extract longer and more effective contexts. To address this issue, we introduce a novel conversational ASR system, extending the Conformer encoder-decoder model with cross-modal conversational representation. Our approach leverages a cross-modal extractor that combines pre-trained speech and text models through a specialized encoder and a modal-level mask input. This enables the extraction of richer historical speech context without explicit error propagation. We also incorporate conditional latent variational modules to learn conversational level attributes such as role preference and topic coherence. By introducing both cross-modal and conversational representations into the decoder, our model retains context over longer sentences without information loss, achieving relative accuracy improvements of 8.8% and 23% on Mandarin conversation datasets HKUST and MagicData-RAMC, respectively, compared to the standard Conformer model.
title Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2310.14278