GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Harry, Partridge, Kurt, Zhu, Pai, Chen, Neng, Park, Hyun Jin, Agarwal, Dhruuv, Wang, Quan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914306226913280
author Zhang, Harry
Partridge, Kurt
Zhu, Pai
Chen, Neng
Park, Hyun Jin
Agarwal, Dhruuv
Wang, Quan
author_facet Zhang, Harry
Partridge, Kurt
Zhu, Pai
Chen, Neng
Park, Hyun Jin
Agarwal, Dhruuv
Wang, Quan
contents Spoken Keyword Spotting (KWS) is the task of distinguishing between the presence and absence of a keyword in audio. The accuracy of a KWS model hinges on its ability to correctly classify examples close to the keyword and non-keyword boundary. These boundary examples are often scarce in training data, limiting model performance. In this paper, we propose a method to systematically generate adversarial examples close to the decision boundary by making insertion/deletion/substitution edits on the keyword's graphemes. We evaluate this technique on held-out data for a popular keyword and show that the technique improves AUC on a dataset of synthetic hard negatives by 61% while maintaining quality on positives and ambient negative audio data.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14814
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples
Zhang, Harry
Partridge, Kurt
Zhu, Pai
Chen, Neng
Park, Hyun Jin
Agarwal, Dhruuv
Wang, Quan
Sound
Computation and Language
Audio and Speech Processing
Spoken Keyword Spotting (KWS) is the task of distinguishing between the presence and absence of a keyword in audio. The accuracy of a KWS model hinges on its ability to correctly classify examples close to the keyword and non-keyword boundary. These boundary examples are often scarce in training data, limiting model performance. In this paper, we propose a method to systematically generate adversarial examples close to the decision boundary by making insertion/deletion/substitution edits on the keyword's graphemes. We evaluate this technique on held-out data for a popular keyword and show that the technique improves AUC on a dataset of synthetic hard negatives by 61% while maintaining quality on positives and ambient negative audio data.
title GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2505.14814