Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Everson, Kevin, Gu, Yile, Yang, Huck, Shivakumar, Prashanth Gurunath, Lin, Guan-Ting, Kolehmainen, Jari, Bulyko, Ivan, Gandhe, Ankur, Ghosh, Shalini, Hamza, Wael, Lee, Hung-yi, Rastrow, Ariya, Stolcke, Andreas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909063127760896
author Everson, Kevin
Gu, Yile
Yang, Huck
Shivakumar, Prashanth Gurunath
Lin, Guan-Ting
Kolehmainen, Jari
Bulyko, Ivan
Gandhe, Ankur
Ghosh, Shalini
Hamza, Wael
Lee, Hung-yi
Rastrow, Ariya
Stolcke, Andreas
author_facet Everson, Kevin
Gu, Yile
Yang, Huck
Shivakumar, Prashanth Gurunath
Lin, Guan-Ting
Kolehmainen, Jari
Bulyko, Ivan
Gandhe, Ankur
Ghosh, Shalini
Hamza, Wael
Lee, Hung-yi
Rastrow, Ariya
Stolcke, Andreas
contents In the realm of spoken language understanding (SLU), numerous natural language understanding (NLU) methodologies have been adapted by supplying large language models (LLMs) with transcribed speech instead of conventional written text. In real-world scenarios, prior to input into an LLM, an automated speech recognition (ASR) system generates an output transcript hypothesis, where inherent errors can degrade subsequent SLU tasks. Here we introduce a method that utilizes the ASR system's lattice output instead of relying solely on the top hypothesis, aiming to encapsulate speech ambiguities and enhance SLU outcomes. Our in-context learning experiments, covering spoken question answering and intent classification, underline the LLM's resilience to noisy speech transcripts with the help of word confusion networks from lattices, bridging the SLU performance gap between using the top ASR hypothesis and an oracle upper bound. Additionally, we delve into the LLM's robustness to varying ASR performance conditions and scrutinize the aspects of in-context learning which prove the most influential.
format Preprint
id arxiv_https___arxiv_org_abs_2401_02921
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
Everson, Kevin
Gu, Yile
Yang, Huck
Shivakumar, Prashanth Gurunath
Lin, Guan-Ting
Kolehmainen, Jari
Bulyko, Ivan
Gandhe, Ankur
Ghosh, Shalini
Hamza, Wael
Lee, Hung-yi
Rastrow, Ariya
Stolcke, Andreas
Computation and Language
Audio and Speech Processing
In the realm of spoken language understanding (SLU), numerous natural language understanding (NLU) methodologies have been adapted by supplying large language models (LLMs) with transcribed speech instead of conventional written text. In real-world scenarios, prior to input into an LLM, an automated speech recognition (ASR) system generates an output transcript hypothesis, where inherent errors can degrade subsequent SLU tasks. Here we introduce a method that utilizes the ASR system's lattice output instead of relying solely on the top hypothesis, aiming to encapsulate speech ambiguities and enhance SLU outcomes. Our in-context learning experiments, covering spoken question answering and intent classification, underline the LLM's resilience to noisy speech transcripts with the help of word confusion networks from lattices, bridging the SLU performance gap between using the top ASR hypothesis and an oracle upper bound. Additionally, we delve into the LLM's robustness to varying ASR performance conditions and scrutinize the aspects of in-context learning which prove the most influential.
title Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2401.02921