Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jung, Yeonjoon, Lee, Jaeseong, Choi, Seungtaek, Lee, Dohyeon, Kim, Minsoo, Hwang, Seung-won
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910658041217024
author Jung, Yeonjoon
Lee, Jaeseong
Choi, Seungtaek
Lee, Dohyeon
Kim, Minsoo
Hwang, Seung-won
author_facet Jung, Yeonjoon
Lee, Jaeseong
Choi, Seungtaek
Lee, Dohyeon
Kim, Minsoo
Hwang, Seung-won
contents Recently, pre-trained language models (PLMs) have been increasingly adopted in spoken language understanding (SLU). However, automatic speech recognition (ASR) systems frequently produce inaccurate transcriptions, leading to noisy inputs for SLU models, which can significantly degrade their performance. To address this, our objective is to train SLU models to withstand ASR errors by exposing them to noises commonly observed in ASR systems, referred to as ASR-plausible noises. Speech noise injection (SNI) methods have pursued this objective by introducing ASR-plausible noises, but we argue that these methods are inherently biased towards specific ASR systems, or ASR-specific noises. In this work, we propose a novel and less biased augmentation method of introducing the noises that are plausible to any ASR system, by cutting off the non-causal effect of noises. Experimental results and analyses demonstrate the effectiveness of our proposed methods in enhancing the robustness and generalizability of SLU models against unseen ASR systems by introducing more diverse and plausible ASR noises in advance.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15609
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
Jung, Yeonjoon
Lee, Jaeseong
Choi, Seungtaek
Lee, Dohyeon
Kim, Minsoo
Hwang, Seung-won
Computation and Language
Sound
Audio and Speech Processing
Recently, pre-trained language models (PLMs) have been increasingly adopted in spoken language understanding (SLU). However, automatic speech recognition (ASR) systems frequently produce inaccurate transcriptions, leading to noisy inputs for SLU models, which can significantly degrade their performance. To address this, our objective is to train SLU models to withstand ASR errors by exposing them to noises commonly observed in ASR systems, referred to as ASR-plausible noises. Speech noise injection (SNI) methods have pursued this objective by introducing ASR-plausible noises, but we argue that these methods are inherently biased towards specific ASR systems, or ASR-specific noises. In this work, we propose a novel and less biased augmentation method of introducing the noises that are plausible to any ASR system, by cutting off the non-causal effect of noises. Experimental results and analyses demonstrate the effectiveness of our proposed methods in enhancing the robustness and generalizability of SLU models against unseen ASR systems by introducing more diverse and plausible ASR noises in advance.
title Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.15609