Contrastive and Consistency Learning for Neural Noisy-Channel Model in Spoken Language Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Suyoung, Hwang, Jiyeon, Jung, Ho-Young
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914809723748352
author Kim, Suyoung
Hwang, Jiyeon
Jung, Ho-Young
author_facet Kim, Suyoung
Hwang, Jiyeon
Jung, Ho-Young
contents Recently, deep end-to-end learning has been studied for intent classification in Spoken Language Understanding (SLU). However, end-to-end models require a large amount of speech data with intent labels, and highly optimized models are generally sensitive to the inconsistency between the training and evaluation conditions. Therefore, a natural language understanding approach based on Automatic Speech Recognition (ASR) remains attractive because it can utilize a pre-trained general language model and adapt to the mismatch of the speech input environment. Using this module-based approach, we improve a noisy-channel model to handle transcription inconsistencies caused by ASR errors. We propose a two-stage method, Contrastive and Consistency Learning (CCL), that correlates error patterns between clean and noisy ASR transcripts and emphasizes the consistency of the latent features of the two transcripts. Experiments on four benchmark datasets show that CCL outperforms existing methods and improves the ASR robustness in various noisy environments. Code is available at https://github.com/syoung7388/CCL.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15097
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Contrastive and Consistency Learning for Neural Noisy-Channel Model in Spoken Language Understanding
Kim, Suyoung
Hwang, Jiyeon
Jung, Ho-Young
Computation and Language
Artificial Intelligence
Recently, deep end-to-end learning has been studied for intent classification in Spoken Language Understanding (SLU). However, end-to-end models require a large amount of speech data with intent labels, and highly optimized models are generally sensitive to the inconsistency between the training and evaluation conditions. Therefore, a natural language understanding approach based on Automatic Speech Recognition (ASR) remains attractive because it can utilize a pre-trained general language model and adapt to the mismatch of the speech input environment. Using this module-based approach, we improve a noisy-channel model to handle transcription inconsistencies caused by ASR errors. We propose a two-stage method, Contrastive and Consistency Learning (CCL), that correlates error patterns between clean and noisy ASR transcripts and emphasizes the consistency of the latent features of the two transcripts. Experiments on four benchmark datasets show that CCL outperforms existing methods and improves the ASR robustness in various noisy environments. Code is available at https://github.com/syoung7388/CCL.
title Contrastive and Consistency Learning for Neural Noisy-Channel Model in Spoken Language Understanding
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.15097