Speak & Spell: LLM-Driven Controllable Phonetic Error Augmentation for Robust Dialogue State Tracking

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Jihyun, Im, Solee, Lee, Wonjun, Lee, Gary Geunbae
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909876851048448
author Lee, Jihyun
Im, Solee
Lee, Wonjun
Lee, Gary Geunbae
author_facet Lee, Jihyun
Im, Solee
Lee, Wonjun
Lee, Gary Geunbae
contents Dialogue State Tracking (DST) is a key part of task-oriented dialogue systems, identifying important information in conversations. However, its accuracy drops significantly in spoken dialogue environments due to named entity errors from Automatic Speech Recognition (ASR) systems. We introduce a simple yet effective data augmentation method that targets those entities to improve the robustness of DST model. Our novel method can control the placement of errors using keyword-highlighted prompts while introducing phonetically similar errors. As a result, our method generated sufficient error patterns on keywords, leading to improved accuracy in noised and low-accuracy ASR environments.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06263
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Speak & Spell: LLM-Driven Controllable Phonetic Error Augmentation for Robust Dialogue State Tracking
Lee, Jihyun
Im, Solee
Lee, Wonjun
Lee, Gary Geunbae
Computation and Language
Artificial Intelligence
Dialogue State Tracking (DST) is a key part of task-oriented dialogue systems, identifying important information in conversations. However, its accuracy drops significantly in spoken dialogue environments due to named entity errors from Automatic Speech Recognition (ASR) systems. We introduce a simple yet effective data augmentation method that targets those entities to improve the robustness of DST model. Our novel method can control the placement of errors using keyword-highlighted prompts while introducing phonetically similar errors. As a result, our method generated sufficient error patterns on keywords, leading to improved accuracy in noised and low-accuracy ASR environments.
title Speak & Spell: LLM-Driven Controllable Phonetic Error Augmentation for Robust Dialogue State Tracking
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2409.06263